Dynamically adapting on-device models of packet assistant devices for collaborative processing of assistant requests
By dynamically adapting models and processing roles in the assistant device group, based on individual processing capabilities and collaboration, the problem of insufficient assistant device processing capabilities is solved, higher robustness and accuracy are achieved, latency and data transmission are reduced, and security is improved.
Patent Information
- Application Number
- CN202510670174.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2020-11-13
- Filing Date
- 2020-12-11
- Publication Date
- 2025-09-26
AI Technical Summary
The limited processing power of the assistant device may result in locally implemented components being less robust and/or accurate, particularly on older and/or lower-cost devices that lack the processing power and memory capacity to execute the various components and/or store the associated models.
Dynamically adapt the models and processing roles on the assistant device, based on individual processing capabilities and collaboration with other devices in the group, and implement on-device pipelines across multiple assistant devices in a distributed manner to collaboratively process assistant requests.
It improves the robustness and accuracy of pipelines on devices, reduces latency, reduces network usage frequency and data transmission volume, and enhances user data security.
Smart Images

Figure CN120704633A_ABST
Abstract
Description
Description of the case
[0001] This application is a divisional application of Chinese invention patent application No. 202080100916.8, with the application date of December 11, 2020. Technical Field
[0002] The present disclosure relates to dynamically adapting on-device models of grouped assistant devices for collaborative processing of assistant requests.
[0003] A person can engage in a human-computer conversation with an interactive software application, referred to herein as an "automated assistant" (also referred to as a "chatbot," "interactive personal assistant," "intelligent personal assistant," "personal voice assistant," "conversational agent," etc.). For example, a person (who may be referred to as a "user" when interacting with an automated assistant) may provide commands and / or requests to the automated assistant using spoken natural language input (i.e., spoken utterances), which in some cases may be converted to text and then processed. Commands and / or requests may additionally or alternatively be provided via one or more other input modalities, such as textual (e.g., typed) natural language input, touchscreen input, and / or touchless gesture input (e.g., detected by a camera of a corresponding assistant device). The automated assistant typically responds to commands or requests by providing responsive user interface output (e.g., audible and / or visual user interface output), controlling a smart device, and / or performing other actions.
[0004] Automated assistants often rely on a pipeline of components when processing user requests. For example, a wake-word detection engine can be used to process audio data while monitoring for the occurrence of a spoken wake-word (e.g., "OK Assistant") and, in response to detecting this occurrence, cause processing by other components to occur. As another example, an automatic speech recognition (ASR) engine can be used to process audio data including spoken utterances to generate a transcription of the user's utterance (i.e., a sequence of lexical items and / or other tokens). The ASR engine can process the audio data based on subsequent occurrences of the spoken wake-word detected by the wake-word detection engine and / or in response to other invocations of the automated assistant. As another example, a natural language understanding (NLU) engine can be used to process the text of a request (e.g., text converted from the spoken utterance using ASR) to generate a symbolic representation or belief state that is a semantic representation of the text. For example, a belief state can include an intent corresponding to the text and, optionally, parameters (e.g., slot values) for the intent. Once fully formed (e.g., all mandatory parameters have been resolved) through one or more dialog turns, the belief state represents the action to be performed in response to the spoken utterance. A separate fulfillment component can then utilize the fully formed belief state to perform actions corresponding to the belief state.
[0005] When interacting with an automated assistant, a user utilizes one or more assistant devices (client devices having an automated assistant interface). The component pipeline used to process requests provided at the assistant device may include components executed locally at the assistant device and / or components implemented at one or more remote servers in network communication with the assistant device.
[0006] Efforts have been made to increase the number of components executed locally on assistant devices and / or to improve the robustness and / or accuracy of such components. These efforts are motivated by considerations such as reducing latency, improving data security, reducing network usage, and / or seeking to achieve other technical benefits. As an example, some assistant devices may include a local wake word engine and / or a local ASR engine.
[0007] However, due to the limited processing power of various assistant devices, components implemented locally on the assistant device may be less robust and / or accurate than their cloud-based counterparts. This may be particularly true for older and / or lower-cost assistant devices, which may lack: (a) the processing power and / or memory capacity to execute various components and / or utilize their associated models; and (b) the disk space capacity to store various associated models. Summary of the Invention
[0008] Embodiments disclosed herein relate to dynamically adapting which assistant on-device model(s) are stored locally at assistant devices in an assistant device group and / or adapting assistant processing roles of assistant devices in the assistant device group. In some of these embodiments, for each assistant device in the group, a corresponding on-device model and / or a corresponding processing role are determined based on a collective consideration of the individual processing capabilities of the assistant devices in the group. For example, the on-device model and / or processing role of a given assistant device can be determined based on the individual processing capabilities of the given assistant device (e.g., whether the given assistant device can store those on-device models and perform those processing roles given processor, memory, and / or storage constraints) and given the corresponding processing capabilities of other assistant devices in the group (e.g., whether the other devices can store other necessary on-device models and / or perform other necessary processing roles). In some embodiments, usage data can also be used to determine a corresponding on-device model and / or a corresponding processing role for each assistant device in the group.
[0009] Implementations disclosed herein additionally or alternatively relate to collaboratively utilizing an assistant device group and its associated post-adaptation on-device models and / or post-adaptation processing roles in collaboratively processing an assistant request directed to any one of the assistant devices in the group.
[0010] Through these and other approaches, given the processing capabilities of these assistant devices, on-device models and on-device processing roles can be distributed across multiple different assistant devices in a group. Furthermore, when distributed across multiple assistant devices in a group, the collective robustness and / or capabilities of the on-device models and / or on-device processing roles exceed the robustness and / or capabilities that any one of the assistant devices might possess individually. In other words, the embodiments disclosed herein can effectively implement an on-device pipeline of assistant components distributed across the assistant devices in a group. The robustness and / or accuracy of this distributed pipeline far exceeds the robustness and / or accuracy capabilities of any pipeline that would otherwise be implemented on a single assistant device in the group. The improved robustness and / or accuracy disclosed herein can reduce latency for a greater number of assistant requests. Furthermore, the improved robustness and / or accuracy results in less (or even no) data being transmitted to the remote automated assistant component to resolve assistant requests. This directly results in increased security for user data, reduced network usage, reduced amounts of data transmitted over the network, and / or reduced latency in resolving assistant requests (e.g., locally resolved assistant requests can be resolved with less latency than would be possible with a remote assistant component).
[0011] Certain adaptations may cause one or more assistant devices in the group to lack the engines and / or models necessary for the assistant device to handle many assistant requests on its own. For example, certain adaptations may cause the assistant device to lack any wake word detection capabilities and / or ASR capabilities. However, when in a group and adapted according to the embodiments disclosed herein, the assistant device can collaborate with other assistant devices in the group to collaboratively process assistant requests, with each other assistant device performing its own processing role and utilizing its own on-device models in doing so. Thus, assistant requests directed to the assistant device can still be processed collaboratively with the other assistant devices in the group.
[0012] In various embodiments, adaptation of assistant devices in a group is performed in response to generating the group or modifying the group (e.g., adding or removing assistant devices from the group). As described herein, a group can be generated based on explicit user input indicating a desire for the group and / or automatically generated based on, for example, determining that the assistant devices in the group satisfy proximity conditions relative to each other. In embodiments where adaptation is performed only when a group is created in this manner, the occurrence of two assistant requests received simultaneously at two separate devices in the group (and potentially unable to be processed collaboratively in parallel) can be mitigated. For example, when proximity conditions are taken into account when generating a group, two unrelated simultaneous requests are less likely to be received at two different assistant devices in the group. For example, this is less likely when the assistant devices in the group are all in the same room than when the assistant devices in the group are spread across multiple floors of a house. As another example, when user input explicitly indicates that a group should be generated, non-overlapping assistant requests are more likely to be provided to the assistant devices in the group.
[0013] As a specific example of various embodiments, assume that an assistant device group including a first assistant device and a second assistant device is generated. Further assume that, when the assistant device group is generated, the first assistant device includes a wake-up word engine and a corresponding wake-up word model, a warm cue engine and a corresponding warm cue model, an authentication engine and a corresponding authentication model, and a corresponding on-device ASR engine and a corresponding ASR model. Further assume that, when the assistant device group is generated, the second assistant device also includes the same engine and model (or a variation thereof) as the first assistant device, and additionally includes: an on-device NLU engine and a corresponding NLU model, an on-device fulfillment and a corresponding fulfillment model, and an on-device TTS engine and a corresponding TTS model.
[0014] In response to the first assistant device and the second assistant device being grouped, the assistant device models stored locally at the first assistant device and the second assistant device and / or the corresponding processing roles of the first assistant device and the second assistant device can be adapted. For each of the first assistant device and the second assistant device, the on-device models and processing roles can be determined based on a consideration of the first processing capabilities of the first assistant device and the second assistant device. For example, a set of on-device models can be determined that includes a first subset that can be stored and utilized by the corresponding processing role / engine on the first assistant device. Further, the set can include a second subset that can be stored and utilized by the corresponding processing role on the second assistant device. For example, the first subset can include only ASR models, but the ASR models of the first subset can be more robust and / or accurate than the pre-adaptation ASR models of the first assistant device. Furthermore, they may require more computing resources when performing ASR. However, the first assistant device may only have the ASR models and ASR engines of the first subset, and the pre-adaptation models and engines can be cleared, thereby freeing up computing resources for use when performing ASR using the ASR models of the first subset. Continuing with the example, the second subset can include the same models as the second assistant device previously included, except that the ASR model can be omitted and a more robust and / or more accurate NLU model can replace the pre-adaptation NLU model. The more robust and / or more accurate NLU model may require more resources than the pre-adaptation NLU model, but these resources can be freed up by clearing the pre-adaptation ASR model (and omitting any ASR models from the second subset).
[0015] Then, in collaboratively processing an assistant request directed to any one of the assistant devices in the group, the first assistant device and the second assistant device can collaboratively utilize their associated post-adaptation on-device models and / or post-adaptation processing roles. For example, assume the spoken utterance "OK Assistant, increase the temperature two degrees." The wake-up cue engine of the second assistant device can detect the occurrence of the wake-up cue "OK Assistant." In response, the wake-up cue engine can transmit the command to the first assistant device so that the ASR engine of the first assistant device performs speech recognition on the audio data captured after the wake-up cue. The transcription generated by the ASR engine of the first assistant device can be transmitted to the second assistant device, and the NLU engine of the second assistant device can perform NLU on the transcription. The results of the NLU can be passed to the fulfillment engine of the second assistant device, which can use these NLU results to determine the command to transmit to the smart thermostat to cause it to increase the temperature by two degrees.
[0016] The foregoing is provided as an overview of only some embodiments. These and other embodiments are disclosed in greater detail herein.
[0017] Additionally, some embodiments may include a system comprising one or more user devices, each user device having one or more processors and a memory operatively coupled to the one or more processors, wherein the memory of the one or more user devices stores instructions that, in response to execution of the instructions by the one or more processors of the one or more user devices, cause the one or more processors to perform any of the methods described herein. Some embodiments may also include at least one non-transitory computer-readable medium comprising instructions that, in response to execution of the instructions by the one or more processors, cause the one or more processors to perform any of the methods described herein.
[0018] It should be appreciated that all combinations of the foregoing concepts and additional concepts described in greater detail herein are contemplated as being part of the subject matter disclosed herein. For example, all combinations of the claims appended to this disclosure are contemplated as being part of the subject matter disclosed herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1A is a block diagram of an example assistant ecosystem in which no assistant devices are grouped and adapted, according to implementations disclosed herein.
[0020] Figure 1B1 、 1B2 and 1B3 respectively illustrate Figure 1A An example assistant ecosystem in which a first assistant device and a second assistant device have been grouped and have different examples of which adaptations can be implemented.
[0021] Figure 1C Pictured Figure 1A An example assistant ecosystem is provided in which a first assistant device, a second assistant device, and a third assistant device have been grouped and have examples of adaptations that can be implemented.
[0022] Figure 1D Pictured Figure 1A An example assistant ecosystem is shown in which a third assistant device and a fourth assistant device have been grouped, with examples of adaptations that can be implemented.
[0023] Figure 2 is a flow chart illustrating an example method of adapting on-device models and / or processing roles for assistant devices in a group.
[0024] Figure 3 is a flow chart illustrating an example method that may be implemented by each of a plurality of assistant devices in a group when adapting on-device models and / or processing roles of the assistant devices in the group.
[0025] Figure 4 An example architecture for a computing device is illustrated. DETAILED DESCRIPTION
[0026] Many users may engage with the automated assistant using any of a number of assistant devices. For example, some users may manage a coordinated "ecosystem" of assistant devices that may receive user input directed to the automated assistant and / or may be controlled by the automated assistant, such as one or more smartphones, one or more tablet computers, one or more vehicle computing systems, one or more wearable computing devices, one or more smart televisions, one or more interactive standalone speakers, one or more interactive standalone speakers with displays, one or more IoT devices, and other assistant devices.
[0027] A user can use any of these assistant devices to engage in a human-computer conversation with the automated assistant (assuming the automated assistant client is installed and the assistant device is capable of receiving input). In some cases, these assistant devices can be dispersed around the user's primary residence, secondary residence, workplace, and / or other structures. For example, mobile assistant devices (such as smartphones, tablets, smartwatches, etc.) can be on the user's person and / or where the user last placed them. Other assistant devices (such as traditional desktop computers, smart TVs, interactive standalone speakers, and IoT devices) can be more stationary, however, and may be located in various locations (e.g., rooms) within the user's home or workplace.
[0028] Initially go to Figure 1A, an example assistant ecosystem is illustrated. The example assistant ecosystem includes a first assistant device 110A, a second assistant device 110B, a third assistant device 110C, and a fourth assistant device 110D. Assistant devices 110A through 110D can all be located within a home, business, or other environment. Furthermore, assistant devices 110A through 110D can all be linked together in one or more data structures or otherwise associated with one another. For example, four assistant devices 110A through 110D can all be registered with the same user account, the same set of user accounts, registered with a particular structure, and / or assigned to a particular structure in a device topology representation. For each of assistant devices 110A through 110D, the device topology representation can include a corresponding unique identifier and can optionally include corresponding unique identifiers for other devices that are not assistant devices (but can interact via the assistant devices), such as IoT devices that do not include an assistant interface. Furthermore, the device topology representation can specify device attributes associated with the respective assistant devices 110A through 110D. Device properties for a given assistant device may indicate, for example, one or more input and / or output modalities supported by the respective assistant device, the processing capabilities of the respective assistant device, the brand, model, and / or unique identifier (e.g., serial number) of the respective assistant device (based on which the processing capabilities may be determined), and / or other properties. As another example, four assistant devices may all be linked together, or otherwise associated with one another, based on being connected to the same wireless network (such as a secure access wireless network) and / or based on collectively communicating peer-to-peer with one another (e.g., via Bluetooth and after pairing). In other words, in some embodiments, multiple assistant devices may be considered linked together, and potentially adapted according to embodiments disclosed herein, based on the multiple assistant devices having a secure network connection with one another and not necessarily being associated with one another in any data structure.
[0029] As a non-limiting working example, first assistant device 110A may be a first type of assistant device, such as a specific model of an interactive standalone speaker with a display and a camera. Second assistant device 110B may be a second type of assistant device, such as a first model of an interactive standalone speaker without a display or camera. Assistant devices 110C and 110D may each be a third type of assistant device, such as a third model of an interactive standalone speaker without a display. The third type (assistant devices 110C and 110D) may have less processing power than the second type (assistant device 110D). For example, the processor of the third type may have less processing power than the processor of the second type. For example, the processor of the third type may lack any GPU, while the processor of the first type includes a GPU. Furthermore, for example, the processor of the third type may have a smaller cache and / or a lower operating frequency than the processor of the second type. As another example, the size of the on-device memory of the third type may be smaller than the size of the on-device memory of the second device (e.g., 1 GB compared to 2 GB). As yet another example, the available disk space of the third type may be smaller than the available disk space of the first type. The available disk space may be different from the currently available disk space. For example, available disk space can be determined as the currently available disk space plus the disk space currently occupied by models on one or more devices. As another example, available disk space can be the total disk space minus any space occupied by the operating system and / or other specific software. Continuing with this working example, the first type and the second type can have the same processing power.
[0030] In addition to being linked together in a data structure, two or more (e.g., all) of assistant devices 110A-110D also selectively communicate with each other, in part, via local area network (LAN) 108. LAN 108 can include other communication topologies such as a wireless network utilizing Wi-Fi, a direct peer-to-peer network such as utilizing Bluetooth, and / or utilizing other communication protocols.
[0031] Assistant device 110A includes assistant client 120A, which may be a standalone application on top of an operating system or may form all or part of the operating system of assistant device 110A. Figure 1A, assistant client 120A includes a wakeup / invocation engine 121A1 and one or more associated on-device wakeup / invocation models 131A1. Wakeup / invocation engine 121A1 can monitor for the occurrence of one or more wakeup or invocation cues and, in response to detecting one or more of the cues, can invoke one or more previously inactive functions of assistant client 120A. For example, invoking assistant client 120A can include causing ASR engine 122A1, NLU engine 123A1, and / or other engines to be activated. For example, it can cause ASR engine 122A1 to process additional audio data frames following the wakeup or invocation cue (where no additional processing of the audio data frames occurred prior to the invocation) and / or can cause assistant client 120A to transmit additional audio data frames and / or other data to be transmitted to cloud-based assistant component 140 for processing (e.g., processing of the audio data frames by a remote ASR engine of cloud-based assistant component 140).
[0032] In some embodiments, wakeup cue engine 121A may continuously process (e.g., if not in "inactive" mode) a stream of audio data frames based on the output of one or more microphones of client device 110A to monitor for the occurrence of a spoken wake word or invocation phrase (e.g., "OK Assistant," "Hey Assistant"). This processing may be performed by wakeup cue engine 121A using one or more wakeup models 131A1. For example, one of wakeup models 131A1 may be a neural network model trained to process audio data frames and generate an output indicating whether one or more wakeup words are present in the audio data. While monitoring for the occurrence of a wakeup word, wakeup cue engine 121A discards (e.g., after temporarily storing in a buffer) any audio data frames that do not include the wakeup word. In addition to or in lieu of monitoring for the occurrence of a wakeup word, wakeup cue engine 121A1 may monitor for the occurrence of other invocation cues. For example, wakeup cue engine 121A1 may also monitor for the press of an invocation hardware button and / or an invocation software button. As another example, and continuing with the working example, when assistant device 110A includes a camera, wakeup cue engine 121A1 can also optionally process image frames from the camera to monitor for the occurrence of an invocation gesture (such as a hand wave while the user's gaze is directed toward the camera) and / or other invocation cues (such as the user's gaze being directed toward the camera together with an indication that the user is speaking).
[0033] exist Figure 1A, assistant client 120A also includes an automatic speech recognition (ASR) engine 122A1 and one or more associated on-device ASR models 132A1. The ASR engine 122A1 can be used to process audio data including spoken utterances to generate a transcription of the user utterance (i.e., a sequence of word terms and / or other tokens). The ASR engine 122A1 can utilize the on-device ASR model 132A1 to process the audio data. The on-device ASR model 132A1 can include, for example, a two-pass ASR model, which is a neural network model and is used by the ASR engine 122A1 to generate a sequence of probabilities over tokens (and the probabilities are used to generate the transcription). As another example, the on-device ASR model 132A1 can include an acoustic model that is a neural network model and a language model that includes a mapping of phoneme sequences to words. The ASR model 122A1 can process the audio data using the acoustic model to generate a phoneme sequence and use the language model to map the phoneme sequence to a specific word term. Additional or alternative ASR models can be utilized.
[0034] exist Figure 1AIn the example, assistant client 120A also includes a natural language understanding (NLU) engine 123A1 and one or more associated on-device NLU models 133A1. NLU engine 123A1 can generate a symbolic representation, or belief state, that is a semantic representation of natural language text, such as transcribed text generated by ASR engine 122A1 or typed text (e.g., typed using the virtual keyboard of assistant device 110A). For example, a belief state can include an intent corresponding to the text and, optionally, parameters (e.g., slot values) of the intent. Once fully formed (e.g., all mandatory parameters have been resolved) through one or more conversation turns, the belief state represents the action to be performed in response to the spoken utterance. In generating the symbolic representation, NLU engine 123A1 can utilize one or more on-device NLU models 133A1. NLU models 133A1 can include one or more neural network models trained to process text and generate output indicating the intent expressed by the text and / or which portion(s) of the text correspond to which parameter(s) of the intent. The NLU model may additionally or alternatively include one or more models that include mappings of text and / or templates to corresponding symbolic representations. For example, a mapping may include mapping the text "what time is it" to the intent "current time" with the parameter "current location." As another example, a mapping may include mapping the template "add [item(s)] to my shopping list" to the intent "insert in shopping list," which has parameters for the items included in the actual natural language corresponding to the [item(s)] in the template.
[0035] exist Figure 1AAssistant client 120A also includes a fulfillment engine 124A1 and one or more associated on-device fulfillment models 134A1. Fulfillment engine 124A1 can utilize the fully-formed symbolic representation from NLU engine 123A1 to perform or cause the performance of an action corresponding to the symbolic representation. The action can include providing a responsive user interface output (e.g., audible and / or visual user interface output), controlling a smart device, and / or performing other actions. When performing or causing the performance of an action, fulfillment engine 124A1 can utilize fulfillment model 134A1. As an example, for the intent "turn on" with a parameter specifying a specific smart light, fulfillment engine 124A1 can utilize fulfillment model 134A1 to identify the network address of the specific smart light and / or the command to be transmitted to transition the specific smart light to the "on" state. As another example, for the intent "current" with a parameter "current location," fulfillment engine 124A1 can utilize fulfillment model 134A1 to identify that the current time at client device 110A should be retrieved and audibly rendered (using TTS engine 125A1).
[0036] exist Figure 1A , assistant client 120A also includes a text-to-speech (TTS) engine 125A1 and one or more associated on-device TTS models 135A1. TTS engine 125A1 can utilize on-device TTS models 135A1 to process text (or its phonetic representation) to generate synthesized speech. The synthesized speech can be audibly rendered via a speaker of the assistant device 110A's local text-to-speech ("TTS") engine (which converts text to speech). The synthesized speech can be generated and rendered as all or part of a response from the automated assistant and / or when prompting a user to define and / or clarify parameters and / or intent (e.g., orchestrated by NLU engine 123A1 and / or a separate dialog state engine).
[0037] exist Figure 1A, assistant client 120A also includes an authentication engine 126A1 and one or more associated on-device authentication models 136A1. Authentication engine 126A1 can utilize one or more authentication techniques to verify which of multiple registered users is interacting with assistant device 110, or if only a single user is registered with assistant device 110, whether it is the registered user (or alternatively a guest / unregistered user) interacting with assistant device 110. As an example, with the permission of the associated users, a text-dependent speaker verification (TD-SV) can be generated and stored for each of the registered users (e.g., in association with their corresponding user profile). Authentication engine 126A1 can utilize the TD-SV model in on-device authentication model 136A1 to generate a corresponding TD-SV and / or process a corresponding portion of the audio data TD-SV to generate a corresponding current TD-SV that can then be compared with the stored TD-SV to determine whether there is a match. As other examples, the authentication engine 126A1 may additionally or alternatively utilize text-independent speaker verification (TI-SV) technology, speaker verification technology, facial verification technology, and / or other verification technology (e.g., PIN entry) and utilize a corresponding on-device authentication model 136A1 to authenticate a particular user.
[0038] exist Figure 1A , the assistant client 120A also includes a warm cue engine 127A1 and one or more associated on-device warm cue models 137A1. The warm cue engine 127A1 can at least selectively monitor the occurrence of one or more warm words or other warm cues, and in response to detecting one or more of the warm cues, cause a specific action to be performed by the assistant client 120A. The warm cues can be in addition to any wake-up words or other wake-up cues, and each of the warm cues can be at least selectively active. Notably, detecting the occurrence of a warm cue causes the specific action to be performed even when the detected occurrence is not preceded by any wake-up cue. Therefore, when the warm cue is one or more specific words, the user can simply say the words without providing any wake-up cues and cause the corresponding specific action to be performed.
[0039] As an example, a "Stop" warm cue may be active at least while a timer or alarm is audibly rendered at assistant device 110A via automated assistant 120A. For example, at such a time, warm cue engine 127A may continue (or at least while VAD engine 128A1 detects voice activity) processing a stream of audio data frames based on output from one or more microphones of client device 110A to monitor for the occurrence of "Stop," "Abort," or another limited set of specific warm words. This processing may be performed by warm cue engine 127A using one of warm cue models 137A1, such as a neural network model trained to process audio data frames and generate an output indicating whether a spoken occurrence of "Stop" is present in the audio data. In response to detecting the occurrence of "Stop," warm cue engine 127A may cause a command to clear the audible timer or alarm. At such a time, warm cue engine 127A may continue (or at least while a presence sensor detects presence) processing an image stream from assistant device 110A's camera to monitor for the occurrence of a hand in a "Stop" gesture. The processing can be performed by the warm cue engine 127A using one of the warm cue models 137A1, such as a neural network model trained to process a frame of visual data and generate an output indicating whether the hand is present and in a "stop" gesture. In response to detecting the presence of the "stop" gesture, the warm cue engine 127A can cause a command to clear an audible sound timer or alarm to be implemented.
[0040] As another example, a "volume up," "volume down," or "next" warm cue may be active at least while music is being audibly rendered at assistant device 110A via automated assistant 120A. For example, at such a time, warm cue engine 127A may continue to process a stream of audio data frames based on the output of one or more microphones of client device 110A. Processing may include monitoring for the occurrence of "volume up" using a first warm cue model in warm cue models 137A1, monitoring for the occurrence of "volume down" using a second warm cue model in warm cue models 137A1, and monitoring for the occurrence of "next" using a third warm cue model in warm cue models 137A1. In response to detecting the occurrence of "volume up," warm cue engine 127A may cause a command to increase the volume of the music to be rendered to be implemented; in response to detecting the occurrence of "volume down," warm cue engine 127A may cause a command to decrease the volume of the music to be rendered to be implemented; and in response to detecting the occurrence of "volume down," warm cue engine 127A may cause a command to be implemented that causes the next music track to be rendered instead of the current music track to be rendered to be implemented.
[0041] exist Figure 1A, assistant client 120A also includes a voice activity detector (VAD) engine 128A1 and one or more associated on-device VAD models 138A1. VAD engine 128A1 can at least selectively monitor audio data for the presence of voice activity and, in response to detecting the presence, cause one or more functions to be performed by assistant client 120A. For example, in response to detecting voice activity, VAD engine 128A1 can cause warm cue engine 121A1 to be activated. As another example, VAD engine 128A1 can be used in a continuous listening mode to monitor audio data for the presence of voice activity and, in response to detecting the presence, cause ASR engine 122A1 to be activated. VAD engine 128A1 can process the audio data using VAD model 138A1 to determine whether voice activity is present in the audio data.
[0042] Specific engines and corresponding models have been described with respect to assistant client 120A. However, it is noted that some engines may be omitted and / or additional engines may be included. It is also noted that, through its various on-device engines and corresponding models, assistant client 120A can fully process many assistant requests, including many assistant requests that are provided as spoken utterances. However, because client device 110A is relatively limited in processing power, there are still many assistant requests that cannot be fully processed locally at assistant device 110A. For example, NLU engine 123A1 and / or corresponding NLU model 133A1 may only cover a subset of all available intents and / or parameters available via the automated assistant. As another example, fulfillment engine 124A1 and / or corresponding fulfillment model may only cover a subset of available fulfillments. As yet another example, ASR engine 122A1 and corresponding ASR model 132A1 may not be robust and / or accurate enough to correctly transcribe various spoken utterances.
[0043] In light of these and other considerations, a cloud-based assistant component 140 may still, at least selectively, be used to perform at least some processing of assistant requests received at assistant device 110A. The cloud-based automated assistant component 140 may include corresponding (and / or additional or alternative) engines and / or models for these assistant devices 110A. However, because the cloud-based automated assistant component 140 can leverage the virtually unlimited resources of the cloud, one or more of its cloud-based counterparts may be more robust and / or accurate with these assistant clients 120A. As an example, in response to a spoken utterance seeking to perform an assistant action not supported by the local NLU engine 123A1 and / or local fulfillment engine 124A1, assistant client 120A may transmit the audio data of the spoken utterance and / or its transcription generated by the ASR engine 122A1 to the cloud-based automated assistant component 140. The cloud-based automated assistant component 140 (e.g., its NLU engine and / or fulfillment engine) may perform more robust processing of such data to support parsing and / or execution of assistant actions. Data is transmitted to the cloud-based automated assistant component 140 via one or more wide area networks (WANs) 109 , such as the Internet or a private WAN.
[0044] Second assistant device 110B includes assistant client 120B, which can be a standalone application on top of an operating system or can form all or part of the operating system of assistant device 110B. Like assistant client 120A, assistant client 120B includes: a wakeup / invocation engine 121B1 and one or more associated on-device wakeup / invocation models 131B1; an ASR engine 122B1 and one or more associated on-device ASR models 132B1; an NLU engine 123B1 and one or more associated on-device NLU models 133B1; a fulfillment engine 124B1 and one or more associated on-device fulfillment models 134B1; a TTS engine 125B1 and one or more associated on-device TTS models 135B1; an authentication engine 126B1 and one or more associated on-device authentication models 136B1; a warm cue engine 127B1 and one or more associated on-device warm cue models 137B1; and a VAD engine 128B1 and one or more associated on-device VAD models 138B1.
[0045] Some or all of the engines and / or models of assistant client 120B may be the same as those of assistant client 120A and / or some or all of the engines and / or models may be different. For example, wakeup cue engine 121B1 may lack functionality to detect wakeup cues in images and / or wakeup model 131B1 may lack a model to process images to detect wakeup cues, while wakeup cue engine 121A1 includes such functionality and wakeup model 131B1 includes such a model. For example, this may be due to assistant device 110A including a camera and assistant device 110B not including a camera. As another example, ASR model 131B1 utilized by ASR engine 122B1 may be different from ASR model 131A1 utilized by ASR engine 122A1. For example, this may be due to different models being optimized for different processor and / or memory capabilities in assistant device 110A and assistant device 110B.
[0046] Specific engines and corresponding models have been described with respect to assistant client 120B. However, it is noted that some engines may be omitted and / or additional engines may be included. It is also noted that, through its various on-device engines and corresponding models, assistant client 120B can fully process many assistant requests, including many assistant requests provided as spoken utterances. However, because client device 110B is relatively limited in processing power, there remain many assistant requests that cannot be fully processed locally at assistant device 110B. In view of these and other considerations, cloud-based assistant component 140 may still be at least selectively used to perform at least some processing of assistant requests received at assistant device 110B.
[0047] Third assistant device 110C includes assistant client 120C, which can be a standalone application on top of an operating system or can form all or part of the operating system of assistant device 110C. Like assistant clients 120A and 120B, assistant client 120C includes: a wakeup / invocation engine 121C1 and one or more associated on-device wakeup / invocation models 131C1; an authentication engine 126C1 and one or more associated on-device authentication models 136C1; a warm lead engine 127C1 and one or more associated on-device warm lead models 137C1; and a VAD engine 128C1 and one or more associated on-device VAD models 138C1. Some or all of the engines and / or models of assistant client 120C can be the same as those of assistant client 120A and / or assistant client 120B, and / or some or all of the engines and / or models can be different.
[0048] Note, however, that unlike assistant client 120A and assistant client 120B, assistant client 120C does not include: any ASR engine or associated model; any NLU engine or associated model; any fulfillment engine or associated model; and any TTS engine or associated model. Further, note that, through its various on-device engines and corresponding models, assistant client 120B can fully process only certain assistant requests (i.e., assistant requests that qualify as warm cues detected by warm cue engine 127C1) and cannot process many assistant requests, such as assistant requests that are provided as spoken utterances and do not qualify as warm cues. In light of these and other considerations, cloud-based assistant component 140 may still be at least selectively used to perform at least some processing of assistant requests received at assistant device 110C.
[0049] Fourth assistant device 110D includes assistant client 120D, which can be a standalone application on top of an operating system or can form all or part of the operating system of assistant device 110D. Like assistant clients 120A, 120B, and 120C, assistant client 120D includes: a wakeup / invocation engine 121D1 and one or more associated on-device wakeup / invocation models 131D1; an authentication engine 126D1 and one or more associated on-device authentication models 136D1; a warm lead engine 127D1 and one or more associated on-device warm lead models 137D1; and a VAD engine 128D1 and one or more associated on-device VAD models 138D1. Some or all of the engines and / or models of assistant client 120C can be the same as those of assistant client 120A, assistant client 120B, and / or assistant client 120C, and / or some or all of the engines and / or models can be different.
[0050] Note, however, that unlike assistant client 120A and assistant client 120B, and like assistant client 120C, assistant client 120D does not include: any ASR engine or associated model; any NLU engine or associated model; any fulfillment engine or associated model; and any TTS engine or associated model. Further, note that, through its various on-device engines and corresponding models, assistant client 120D can fully process only certain assistant requests (i.e., assistant requests that qualify as warm cues detected by warm cue engine 127D1), and cannot process many assistant requests, such as assistant requests that are provided as spoken utterances and do not qualify as warm cues. In light of these and other considerations, cloud-based assistant component 140 may still be at least selectively used to perform at least some processing of assistant requests received at assistant device 110D.
[0051] Now go to Figure 1B1 、 1B2, 1B3, 1C, and 1D illustrate different non-limiting examples of assistant device groups and different non-limiting examples of adaptations that can be performed in response to the assistant device group being generated. Through each of the adaptations, the grouped assistant devices can be collectively utilized to process various assistant requests, and through collective utilization, more robust and / or more accurate processing of these various assistant requests can be performed compared to any one of the group's assistant devices performing individually before the adaptation. This results in various technical advantages, such as those described herein.
[0052] exist Figure 1B1 、 1B2 , 1B3, 1C and 1D, relative to Figure 1A , with Figure 1A The engines and models of the assistant clients with the same reference numerals in the example are not adapted. Figure 1B1 、 1B2 and 1B3, the engines and models of assistant client devices 110C and 110D are not adapted because assistant client devices 110C and 110D are not included in Figure 1B1 、 1B2 and 1B3 in group 101B. However, in Figure 1B1 、 1B2 , 1B3, 1C and 1D, with Figure 1A Different reference numbers (i.e., ending with "2", "3", or "4" instead of "1") for the assistant client's engine and model indicate that it has been compared to Figure 1A Furthermore, an engine or model having a reference number ending in "2" in one figure and ending in "3" in another figure means that a different adaptation of the engine or model has been made between the figures. Likewise, an engine or model having a reference number ending in "4" in a figure means that the adaptation of the engine or model in that figure is different from the figure ending in "2" or "3".
[0053] Initially go to Figure 1B1, device group 101B has been established, with assistant devices 110A and 110B included in device group 101B. In some embodiments, device group 101B can be generated in response to user interface input that explicitly indicates a desire to group assistant devices 110A and 110B. As an example, a user can provide the spoken utterance "group [label for assistant device 110A] and [label for assistant device 110B]" to any of assistant devices 110A through 110D. This spoken utterance can be processed by the corresponding assistant device and / or cloud-based assistant component 140 and interpreted as a request to group assistant devices 110A and 110B, and group 101B can be generated in response to this interpretation. As another example, a registered user of assistant devices 110A through 110D can provide touch input at an application that supports configuring settings for assistant devices 110A through 110D. These touch inputs can explicitly specify that assistant devices 110A and 110B are to be grouped, and that group 101B can be generated in response. As yet another example, one of the example techniques described below for automatically generating device group 101B can alternatively be used to determine that device group 101B should be generated, but user input that explicitly approves the generation of device group 101B can be required before device group 101B is generated. For example, a prompt indicating that device group 101B should be generated can be rendered at one or more of assistant devices 110A through 110D, and device group 101B is actually generated only if affirmative user interface input is received in response to the prompt (and optionally, if the user interface input is verified to be from a registered user).
[0054] In some embodiments, device group 101B may instead be automatically generated. In some of these embodiments, a user interface output indicating the generation of device group 101B may be rendered on one or more of assistant devices 110A-110D to inform the corresponding user and / or registered user of the group that the automated generation of device group 101B via user interface input may be overridden. However, when device group 101B is automatically generated, device group 101B is generated and adapted accordingly without first requesting a desired user interface input explicitly indicating the creation of a particular device group 101B (although earlier input may indicate general approval for group creation). In some embodiments, device group 101B may be automatically generated in response to determining that assistant devices 110A and 110B meet one or more proximity conditions relative to each other. For example, the proximity conditions may include assistant devices 110A and 110B being assigned to the same structure (e.g., a specific home, a specific vacation home, a specific office) and / or the same room (e.g., a kitchen, living room, dining room), or other area within the same structure in a device topology. As another example, a proximity condition can include a sensor signal from each of assistant devices 110A and 110B indicating that they are geographically close to each other. For example, if both assistant devices 110A and 110B detect occurrences of the wake word at or near the same time (e.g., within one second of it) consistently (e.g., greater than 70% of the time or other threshold), this can indicate that they are geographically close to each other. Also, for example, one of assistant devices 110A and 110B can transmit a signal (e.g., ultrasound), and the other of assistant devices 110A and 110B can attempt to detect the transmitted signal. If the other of assistant devices 110A and 110B detects the transmitted signal, optionally with a threshold strength, it can indicate that they are geographically close to each other. Additional and / or alternative techniques for determining temporal proximity and / or automatically generating device groups can be utilized.
[0055] Regardless of how group 101B is generated, Figure 1B1One example of adaptations that may be performed on assistant devices 110A and 110B in response to their inclusion in group 101B is shown. In various embodiments, one or both of assistant clients 120A and 120B may determine the adaptations that should be performed and cause those adaptations to occur. In other embodiments, one or more engines of cloud-based assistant component 140 may additionally or alternatively determine the adaptations that should be performed and cause those adaptations to occur. As described herein, the adaptations to be performed may be determined based on considerations of the processing capabilities of both assistant clients 120A and 120B. For example, the adaptations may seek to utilize as much collective processing power as possible while ensuring that the individual processing power of each of the assistant devices is sufficient for engines and / or models to be stored and utilized locally at the assistant. Further, the adaptations to be performed may also be determined based on usage data reflecting metrics related to actual usage of the assistant devices in the group and / or other non-grouped assistant devices of the ecosystem. For example, if processing power allows either a larger but more accurate wake-up cue model or a larger but more robust warm-cue model, but not both, usage data can be used to select between the two options. For example, if the usage data reflects rare (or even no) use of warm words and / or detection of wake words typically barely exceeds a threshold and / or false negatives of wake words are commonly encountered, then a larger but more accurate wake-up cue model can be selected. On the other hand, if the usage data reflects frequent use of warm words and / or detection of wake words consistently exceeds a threshold and / or false negatives of wake words are rare, then a larger but more accurate warm-cue model can be selected. Considering such usage data can be a factor in determining Figure 1B1 、 1B2 The adaptation of 1B3 is still the factor that is chosen, because Figure 1B1 、 1B2 or 1B3 respectively show different adaptations of the same group 101B.
[0056] exist Figure 1B1 , fulfillment engine 124A1 and fulfillment model 134A1, as well as TTS engine 125A1 and TTS model 135A1, have been cleared from assistant device 110A. Further, assistant device 110A has a different ASR engine 122A2 and a different on-device ASR model 132A2, as well as a different NLU engine 123A2 and a different on-device NLU model 133A2. The different engines can be downloaded at assistant device 110A from a local model repository 150 accessible via interaction with cloud-based assistant component 140. Figure 1B1In the example, wakeup cue engine 121B1, ASR engine 122B1, and authentication engine 126B1, along with their corresponding models 131B1, 133B1, and 136B1, have been cleared from assistant device 110B. Furthermore, assistant device 110B now has a different NLU engine 123B2 and a different on-device NLU model 133B2, a different fulfillment engine 124B2 and fulfillment model 133B2, and a different warm word engine 127B2 and warm word model 137B2. The different engines can be downloaded at assistant client 110B from a local model repository 150 accessible via interaction with cloud-based assistant component 140.
[0057] Compared to ASR engine 122A1 and ASR model 132A1, ASR engine 122A2 and ASR model 132A2 of assistant device 110A may be more robust and / or more accurate, but may occupy more disk space, utilize more memory, and / or require more processor resources. For example, ASR model 132A1 may include only a single-pass model, and ASR model 132A2 may include a two-pass model.
[0058] Likewise, NLU engine 123A2 and NLU model 133A2 may be more robust and / or more accurate, but may occupy more disk space, utilize more memory, and / or require more processor resources than NLU engine 123A1 and NLU model 133A1. For example, NLU model 133A1 may include only intents and parameters of the first category, such as "lighting control," but NLU model 133A2 may include intents for "lighting control" as well as "thermostat control," "smart lock control," and "reminders."
[0059] Thus, ASR engine 122A2, ASR model 132A2, NLU engine 123A2, and NLU model 133A2 are improvements over their replaced counterparts. However, it should be noted that the processing power of assistant device 110A may hinder the storage and / or use of ASR engine 122A2, ASR model 132A2, NLU engine 123A2, and NLU model 133A2 without first purging fulfillment engine 124A1 and fulfillment model 134A1, as well as TTS engine 125A1 and TTS model 135A1. Simply purging such models from assistant device 110A without complementary adaptation to and collaborative processing with assistant device 110B will result in assistant client 120A lacking the ability to fully process various assistant requests locally (i.e., without necessarily leveraging one or more cloud-based assistant components 140).
[0060] Thus, assistant device 110B is supplementally adapted, and collaborative processing between assistant devices 110A and 110B occurs after the adaptation. Compared to NLU engine 123B1 and NLU model 133B1, NLU engine 123B2 and NLU model 133B2 of assistant device 110B can be more robust and / or more accurate, but take up more disk space, utilize more memory, and / or require more processor resources. For example, NLU model 133B1 may only include intents and parameters of a first classification, such as "lighting control." However, NLU model 133B2 may cover a larger number of intents and parameters. It is noted that the intents and parameters covered by NLU model 133B2 may be limited to intents that are not already covered by NLU model 133A2 of assistant client 120A. This may prevent duplication of functionality between assistant clients 120A and 120B and expand collective capabilities when assistant clients 120A and 120B collaboratively process assistant requests.
[0061] Likewise, fulfillment engine 124B2 and fulfillment model 134B2 may be more robust and / or more accurate, but may occupy more disk space, utilize more memory, and / or require more processor resources than fulfillment engine 124B1 and fulfillment model 124B1. For example, fulfillment model 124B1 may include only the fulfillment capabilities for a single category of NLU model 133B1, but fulfillment model 124B2 may include the fulfillment capabilities for all categories of NLU model 133B2 as well as NLU model 133A2.
[0062] Thus, fulfillment engine 124B2, fulfillment model 134B2, NLU engine 123B2, and NLU model 133B2 are improvements over their replacement counterparts. However, without first purging the purged models and purged engines from assistant device 110B, the processing power of assistant device 110B may hinder the storage and / or use of fulfillment engine 124B2, fulfillment model 134B2, NLU engine 123B2, and NLU model 133B2. Simply purging such models from assistant device 110B without additional adaptation to and collaborative processing with assistant device 110A will result in assistant client 120B lacking the ability to fully natively process various assistant requests.
[0063] Compared to the warm clue engine 127B1 and warm clue model 127B1, the warm clue engine 127B2 and warm clue model 137B2 of the client device 110B do not take up any additional disk space, utilize more memory, or require more processor resources. For example, they may require the same or even less processing power. However, the warm clue engine 127B2 and warm clue model 137B2 cover warm clues, which are in addition to the warm clues covered by the warm clue engine 127B1 and warm clue model 127B1, and are in addition to the warm clues covered by the warm clue engine 127A1 and warm clue model 127A1 of the assistant client 120A.
[0064] exist Figure 1B1 In a configuration such as , assistant client 120A can be assigned the following processing roles: monitor for wakeup cues, perform ASR, perform NLU on a first set of classifications, perform authentication, monitor a first set of warm cues, and perform VAD. Assistant client 120B can be assigned the following processing roles: perform NLU on a second set of classifications, perform fulfillment, perform TTS, and monitor a second set of warm cues. The processing roles can be passed and stored at each of assistant clients 120A, and coordination of the processing of various assistant requests can be performed by one or both of assistant clients 120A and 120B.
[0065] As use Figure 1B1As an example of collaborative processing of an adapted assistant request, assume that the spoken utterance "OK Assistant, turn on the kitchen lights" is provided, and assistant device 120A is the lead device coordinating the processing. Wake-up cue engine 121A1 of assistant client 120A can detect the presence of the wake-up cue "OK Assistant." In response, wake-up cue engine 121A1 can cause ASR engine 122A2 to process audio data captured following the wake-up cue. Wake-up cue engine 121A1 can also optionally transmit a command locally to assistant device 110B to cause it to transition from a low-power state to a high-power state, thereby preparing assistant client 120B to perform certain processing requested by the assistant. The audio data processed by ASR engine 122A2 can be audio data captured by assistant device 110A's microphone and / or audio data captured by assistant device 110B's microphone. For example, the command transmitted to assistant device 110B to cause it to transition to a high-power state can also cause it to capture audio data locally and optionally transmit this audio data to assistant client 120A. In some embodiments, assistant client 120A can determine whether to use received audio data or locally captured audio data based on an analysis of characteristics of the corresponding instances of audio data. For example, an instance of audio data can be utilized over another instance based on having a lower signal-to-noise ratio and / or capturing spoken utterances at a higher volume.
[0066] The transcription generated by ASR engine 122A2 can be passed to NLU engine 123A2 to perform NLU on the transcription for the first set of classifications, and also transmitted to assistant client 120B so that NLU engine 123B2 can perform NLU on the transcription for the second set of classifications. The results of the NLU performed by NLU engine 123B2 can be transmitted to assistant client 120A, which can determine which results (if any) to use based on these results and the results from NLU engine 123A2. For example, assistant client 120A can utilize the result with the highest probability intent, as long as the probability meets some threshold. For example, the result including the intent "turn on" and parameters specifying the identifier of "kitchen light" can be utilized. Note that if no probability meets the threshold, the NLU engine of cloud-based assistant component 140 can optionally be used to perform NLU. Assistant client 120A can transmit the NLU result with the highest probability to assistant client 120B. Fulfillment engine 124B2 of assistant client 120B can use these NLU results to determine a command to transmit to the "kitchen lights" to turn them to the "on" state and transmit this command over LAN 108. Optionally, fulfillment engine 124B2 can use TTS engine 125B1 to generate synthesized speech confirming the execution of "turn on the kitchen lights." In this case, the synthesized speech can be rendered by assistant client 120B at assistant device 110B and / or transmitted to assistant device 110A for rendering by assistant client 120A.
[0067] As another example of collaborative processing of assistant requests, assume that assistant client 120A is rendering an alarm for a local timer at assistant client 120A that just expired. Further assume that the warm cue monitored by warm cue engine 127B2 includes "stop" and the warm cue monitored by warm cue engine 127A1 does not include "stop". Finally, assume that while the alarm is being rendered, the spoken word "stop" is provided and captured in audio data detected via the microphone of assistant client 120B. Warm word engine 127B2 can process the audio data and determine that the word "stop" appeared. Further, warm word engine 127B2 can determine that the appearance of the word stop is directly mapped to a command to clear an audible sound timer or alarm. The command can be transmitted by assistant client 120B to assistant client 120A to cause assistant client 120A to implement the command and clear the audible sound timer or alarm. In some embodiments, warm word engine 127B2 can monitor for the appearance of "stop" only in certain circumstances. In these embodiments, in response to rendering an alert or in anticipation of rendering an alert, assistant client 120A can transmit a command to cause warm word engine 127B2 to monitor for occurrences of “stop.” The command can cause monitoring to occur within a certain time period, or alternatively until a stop monitoring command is sent.
[0068] Now go to Figure 1B2 , the same group 101B is shown. Figure 1B2 In, with Figure 1B1 The same adaptation has been performed as in , except that the ASR engine 121A1 and the ASR model 132A1 are not replaced by the ASR engine 122A2 and the ASR model 132A2. Instead, the ASR engine 121A1 and the ASR model 132A1 are retained, and an additional ASR engine 122A3 and an additional ASR model 132A3 are provided.
[0069] ASR engine 121A1 and ASR model 132A1 may be used for speech recognition of utterances in a first language (eg, English), and additional ASR engine 122A3 and additional ASR model 132A3 may be used for speech recognition of utterances in a second language (eg, Spanish). Figure 1B1 ASR engine 121A2 and ASR model 132A2 can also be used for the first language and can be more robust and / or accurate than ASR engine 121A1 and ASR model 132A1. However, the processing power of assistant client 120A may prevent ASR engine 121A2 and ASR model 132A2 from being stored locally along with ASR engine 122A3 and additional ASR model 132A3. However, the processing power supports the storage and utilization of ASR engine 121A1 and ASR model 132A1 along with additional ASR engine 122A3 and additional ASR model 132A3.
[0070] exist Figure 1B2 In the example, the locally stored ASR engine 121A1 and the ASR model 132A1 and the additional ASR engine 122A3 and the additional ASR model 132A3 are replaced Figure 1B1 The decisions of ASR engine 121A2 and ASR model 132A2 can be based on usage statistics indicating that the spoken utterances provided at assistant devices 110A and 110B (and / or assistant devices 110C and 110D) include first language spoken utterances and second language spoken utterances. Figure 1B1 In the example, the usage statistics may indicate that only spoken utterances in the first language are present, leading to the selection of Figure 1B1 A more robust ASR engine 121A2 and ASR model 132A2 in.
[0071] Now go to Figure 1B3 , the same group 101B is illustrated again. Figure 1B3 In, with Figure 1B1 The same adaptations have been made in , except that: (1) ASR engine 121A1 and ASR model 132A1 are replaced by ASR engine 122A4 and ASR model 132A4, rather than by ASR engine 122A2 and ASR model 132A2; (2) no warm cue engine or warm cue model exists on assistant device 110B; and (3) ASR engine 122B4 and ASR model 132B4 are stored and utilized locally on assistant device 110B.
[0072] ASR engine 122A4 and ASR model 132A4 can be used to perform the first part of speech recognition, while ASR engine 122B4 and ASR model 132B4 can be used to perform the second part of speech recognition. For example, ASR engine 122A4 can utilize ASR model 132A4 to generate an output, which is transmitted to assistant client 120B, and ASR engine 122B4 can process the output when generating the speech recognition. As a specific example, the output can be a graph representing candidate recognitions, and ASR engine 122B4 can perform a beam search on the graph when generating the speech recognition. As another specific example, ASR model 132A4 can be the initial / downstream part (i.e., the first neural network layer) of an end-to-end speech recognition model, and ASR model 132B4 can be the later / upstream part (i.e., the second neural network layer) of the end-to-end speech recognition model. In this example, the end-to-end model is split between assistant devices 110A and 110B, and the output can be the state of the final layer (e.g., embedding) of the initial part after processing. As another example, the ASR model 132A4 can be an acoustic model, and the ASR model 132B4 can be a language model. In this example, the output can indicate a sequence of phonemes or a probability distribution sequence of phonemes, and the ASR engine 122B4 can use the language model to select a transcription / recognition corresponding to the sequence.
[0073] The robustness and / or accuracy of the ASR engine 122A4, ASR model 132A4, ASR engine 122B4, and ASR model 132B4 working collectively may exceed Figure 1B1 Further, the processing power of assistant client 120A and assistant client 120B may hinder the ASR models 132A4 and 132B4 from being stored and utilized separately on either of the devices. However, the processing power may support splitting the models and splitting the processing roles between the ASR engines 122A4 and 122B4 as described herein. Note that on assistant device 110B, clearing the warm cue engine and warm cue model may support the storage and utilization of the ASR engine 122B4 and ASR model 132B4. In other words, the processing power of assistant device 110B may not support the warm cue engine and warm cue model being stored and utilized separately on either of the devices. Figure 1B3 Other engines and models are shown as being stored and / or utilized together.
[0074] exist Figure 1B3 In the example of FIG. 1 , the locally stored ASR engine 122A4, the ASR model 132A4, the ASR engine 122B4, and the ASR model 132B4 are replaced Figure 1B1Decisions made by ASR engine 121A2 and ASR model 132A2 can be based on usage statistics indicating that speech recognition at assistant devices 110A and 110B (and / or assistant devices 110C and 110D) is often low-confidence and / or often inaccurate. For example, the usage statistics can indicate that a confidence metric for the recognition is below average (e.g., based on an average for a population of users) and / or that the recognition is often corrected by the user (e.g., by editing a display of the transcription).
[0075] Now go to Figure 1C , device group 101C has been established, wherein assistant devices 110A, 110B, and 110C are included in device group 101C. In some embodiments, device group 101C can be generated in response to a user interface input that explicitly indicates a desire to group assistant devices 110A, 110B, and 110C. For example, the user interface input can indicate the creation of device group 101C from scratch or, alternatively, the addition of assistant device 110C to device group 101B ( Figure 1B1 、 1B2 and 1B3) thereby creating a desire to modify group 101C. In some embodiments, device group 101C may be automatically generated as a replacement. For example, device group 101B ( Figure 1B1 、 1B2 and 1B3) may have been previously generated based on a determination that assistant devices 110A and 110B are in close proximity, and after device group 101B is created, assistant device 110C may be moved by the user so that it is now in close proximity to devices 110A and 110B. Consequently, assistant device 110C may be automatically added to device group 101B, thereby creating modified group 101C.
[0076] Regardless of how Group 101C is generated, Figure 1C One example of adaptations that may be made to assistant devices 110A, 110B, and 110C in response to their inclusion in group 101C is shown.
[0077] exist Figure 1C , the assistant device 110B already has Figure 1B3 Further, the assistant device 110A has the same adaptation as Figure 1B3, except that: (1) authentication engine 126A1 and VAD engine 128A1 and their corresponding models 13A1 and 138A1 have been cleared; (2) wakeup cue engine 121A1 and wakeup cue model 131A1 have been replaced with wakeup cue engine 121A2 and wakeup cue model 131A2; and (3) no warm cue engine or warm cue model exists on assistant device 110B. The models and engines stored on assistant device 110C are not adapted. However, assistant client 120C can be adapted to enable collaborative processing of assistant requests with assistant client 120A and assistant 120B.
[0078] exist Figure 1CIn the example, authentication engine 126A1 and VAD engine 128A1 have been cleared from assistant device 110A because their counterparts already exist on assistant device 110C. In some embodiments, authentication engine 126A1 and / or VAD engine 128A1 may be cleared only after some or all data from these components has been merged with their counterparts already existing on assistant device 110C. As an example, authentication engine 126A1 may store voice embeddings for both the first and second users, but authentication engine 126C1 may only store the first user's voice embedding. Before purging authentication engine 126A1, the second user's voice embedding may be locally transmitted to authentication engine 126C1 so that such voice embedding can be utilized by authentication engine 126C1, ensuring that pre-adaptation capabilities are maintained after adaptation. As another example, authentication engine 126A1 may store instances of audio data that separately captured the second user's utterance and were used to generate the second user's voice embedding, and authentication engine 126C1 may lack any voice embeddings for the second user. Prior to clearing authentication engine 126A1, an instance of audio data may be locally transferred from authentication engine 126A1 to authentication engine 126C1 so that the instance of audio data may be utilized by authentication engine 126C1 to generate a sound embedding for the second user, utilizing on-device authentication model 136C1, to ensure that pre-adaptation capabilities are maintained after adaptation. Further, wakeup cue engine 121A1 and wakeup cue model 131A1 have been replaced with a wakeup cue engine 121A2 and wakeup cue model 131A2 of smaller storage size. For example, wakeup cue engine 121A1 and wakeup cue model 131A1 support detection of spoken wakeup cues and image-based wakeup cues, while wakeup cue engine 121A2 and wakeup cue model 131A2 only support detection of image-based wakeup cues. Optionally, before clearing the wakeup cue engine 121A1 and the wakeup cue model 131A1, the personalization, training examples, and / or other settings for the image-based wakeup cue portion of the wakeup cue engine 121A1 and the wakeup cue model 131A1 can be merged with the wakeup cue engine 121A2 and the wakeup cue model 131A2, or otherwise shared with them. The wakeup cue engine 121C1 and the wakeup cue model 131C1 only support detecting spoken wakeup cues. Therefore, the wakeup cue engine 121A2 and the wakeup cue model 131A2, as well as the wakeup cue engine 121C1 and the wakeup cue model 131C1, collectively support detecting spoken and image-based wakeup cues. Optionally, the personalization and / or other settings for the spoken cue portion of the wakeup cue engine 121A1 and the wakeup cue model 131A1 can be transmitted to the client device 110C to be merged with the wakeup cue engine 121C1 and the wakeup cue model 131A1, or otherwise shared with them.
[0079] Furthermore, replacing the wakeup cue engine 121A1 and wakeup cue model 131A1 with a smaller wakeup cue engine 121A2 and wakeup cue model 131A2 provides additional storage space. This additional storage space, along with the additional storage space provided by clearing the authentication engine 126A1 and VAD engine 128A1 and their corresponding models 13A1 and 138A1, provides space for the warm cue engine 127A2 and warm cue model 137A2 (which are collectively larger than the warm cue engine 127A1 and warm cue model they replace). The warm cue engine 127A2 and warm cue model 137A2 can be used to monitor different warm cues than those monitored by the warm cue engine 127C1 and warm cue model 137C1.
[0080] As use Figure 1C As an example of collaborative processing of adapted assistant requests, assume that the spoken utterance "OK Assistant, turn on the kitchen lights" is provided and assistant device 120A is the lead device for the coordinated processing. Wake-up cue engine 121C1 of assistant client 110C can detect the presence of the wake-up cue "OK Assistant." In response, wake-up cue engine 121C1 can transmit a command to assistant devices 110A and 110B causing ASR engines 122A4 and 122B4 to collaboratively process audio data captured following the wake-up cue. The processed audio data can be audio data captured by a microphone of assistant device 110C and / or audio data captured by a microphone of assistant device 110B and / or assistant device 110C.
[0081] The transcription generated by ASR engine 122B4 can be passed to NLU engine 123B2 to perform NLU on the transcription for the second set of classifications and also transmitted to assistant client 120A for NLU engine 123A2 to perform NLU on the transcription for the first set of classifications. The results of the NLU performed by NLU engine 123B2 can be transmitted to assistant client 120A, which can determine which results (if any) to use based on these results and the results from NLU engine 123A2. Assistant client 120A can transmit the NLU results with the highest probability to assistant client 120B. Fulfillment engine 124B2 in assistant client 120B can use these NLU results to determine a command to transmit to the "kitchen lights" to turn them to the "on" state and transmit this command via LAN 108. Optionally, fulfillment engine 124B2 can utilize TTS engine 125B1 to generate synthesized speech confirming the execution of "Turn on the kitchen lights." In this case, the synthesized speech can be rendered at assistant device 110B by assistant client 120B, transmitted to assistant device 110A for rendering by assistant client 120A, and / or transmitted to assistant device 110C for rendering by assistant client 120C.
[0082] Now go to Figure 1D , device group 101D has been established, wherein assistant devices 110C and 110D are included in device group 101D. In some embodiments, device group 101D can be generated in response to a user interface input that explicitly indicates a desire to group assistant devices 110C and 110D. In some embodiments, device group 101D can alternatively be automatically generated.
[0083] Regardless of how group 101D is generated, Figure 1D One example of adaptations that may be made to assistant devices 110C and 110D in response to their inclusion in group 101D is shown.
[0084] exist Figure 1D In the example, the wakeup cue engine 121C1 and wakeup cue model 131C1 of assistant device 110C are replaced with wakeup cue engine 121C2 and wakeup cue model 131C2. Furthermore, the authentication engine 126D1 and authentication model 136D1, as well as the VAD engine 128D1 and VAD model 138D2, are cleared from assistant device 110D. Furthermore, the wakeup cue engine 121D1 and wakeup cue model 131D1 of assistant device 110D are replaced with wakeup cue engine 121D2 and wakeup cue model 131D2, and the warm cue engine 127D1 and warm cue model 137D1 are replaced with warm cue engine 127D2 and warm cue model 137D2.
[0085] The previous wake-up cue engine 121C1 and wake-up cue model 131C1 may be used only to detect a first set of one or more wake-up words, such as "Hey Assistant" and "OK Assistant." On the other hand, the wake-up cue engine 121C2 and wake-up cue model 131C2 may only detect an alternating second set of one or more wake-up words, such as "Hey Computer" and "OK Computer." The previous wake-up cue engine 121D1 and wake-up cue model 131D1 may also be used only to detect the first set of one or more wake-up words, and the wake-up cue engine 121D2 and wake-up cue model 131D2 may also be used only to detect the first set of one or more wake-up words. However, the wake-up cue engine 121D2 and wake-up cue model 131D2 are larger than their replaced counterparts and are also more robust (e.g., more robust to background noise) and / or more accurate. Purging the engine and model from assistant device 110D may enable the utilization of the larger wake-up cue engine 121D2 and wake-up cue model 131D2. Further, collectively, the wake-up cue engine 121C2 and the wake-up cue model 131C2 and the wake-up cue engine 121D2 and the wake-up cue model 131D2 support detecting two sets of wake-up words, while each of the assistant clients 120C and 120D can only detect the first set before adaptation.
[0086] The warm clue engine 127D2 and warm clue model 137D2 of the assistant device 110D may require more computing power than the replaced wake-up clue engine 127D1 and wake-up clue model 137D1. However, these capabilities are available by clearing the engine and model from the assistant device 110D. Moreover, the warm clues monitored by the warm clue engine 127D2 and warm clue model 137D1 may be a supplement to the warm clues monitored by the warm clue engine 127C1 and warm clue model 137D1. Before adaptation, the wake-up clues monitored by the wake-up clue engine 127D1 and wake-up clue model 137D1 are the same as the warm clues monitored by the warm clue engine 127C1 and warm clue model 137D1. Therefore, through collaborative processing, the assistant client 120C and the assistant client 120D can monitor a larger number of wake-up clues.
[0087] It should be noted that in Figure 1DIn the example of , there are many assistant requests that cannot be fully processed by assistant clients 120C and 120D collaboratively on the devices. For example, assistant clients 120C and 120D lack any ASR engines, lack any NLU engines, and lack any fulfillment engines. This may be due to the processing power of assistant devices 110C and 110D not supporting any such engines or models. Therefore, cloud-based assistant component 140 will still need to be used to fully process many assistant requests for spoken utterances of warm cues that are not supported by assistant clients 120C and 120D. However, compared to any processing that would occur individually at the devices before adaptation, Figure 1D The adaptation and the collaborative processing based on the adaptation can still be more robust and / or accurate. For example, the adaptation supports the detection of additional wake-up cues and additional warm cues.
[0088] As an example of collaborative processing that may occur, assume the utterance "OK Computer, play some music" is spoken. In this example, wake-up cue engine 121D2 may detect the wake-up cue "OK Computer". In response, wake-up cue engine 121D2 may cause audio data corresponding to the wake-up cue to be transmitted to assistant client 120C. Authentication engine 126C1 of assistant client 120C may utilize the audio data to determine whether the dictation of the wake-up cue is authenticated to the registered user. Wake-up cue engine 121D2 may also cause the audio data following the spoken utterance to be streamed to cloud-based assistant component 140 for further processing. The audio data may be captured at assistant device 110D or captured at assistant device 110C (e.g., assistant client 120 may transmit a command to assistant client 120C to cause it to capture audio data in response to detection of the wake-up cue by wake-up cue engine 121D2). Further, authentication data based on the output of authentication engine 126C1 may also be transmitted along with the audio data. For example, if the authentication engine 126C1 authenticates the dictation of the wakeup cue to a registered user, the authentication data may include an identifier of the registered user. As another example, if the authentication engine 126C1 does not authenticate the dictation of the wakeup cue to any registered user, the authentication data may include an identifier reflecting that the utterance was provided by a guest user.
[0089] Various specific examples have been referenced Figure 1B1 、 1B2 However, it is noted that various additional or alternative groups may be generated and / or various additional or alternative adaptations may be performed in response to the generation of a group.
[0090] Figure 2is a flowchart illustrating an example method 200 for adapting an on-device model and / or processing role for an assistant device in a group. For convenience, the operations of the flowchart are described with reference to a system performing the operations. The system may include various components of various computer systems, such as Figure 1A 、 Figure 1B1 、 Figure 1B2 、 Figure 1B3 、 Figure 1C and Figure 1D Automation clients 120A to 120D and / or Figure 1A 、 Figure 1B1 、 Figure 1B2 、 Figure 1B3 、 Figure 1C and Figure 1D Furthermore, although the operations of method 200 are shown in a particular order, this is not intended to be limiting. One or more operations may be reordered, omitted, or added.
[0091] In block 252, the system generates the assistant device group. For example, the system may generate the assistant device group in response to a user interface input that explicitly indicates a desire to generate the group. As another example, the system may automatically generate the group in response to determining that one or more conditions are satisfied. As yet another example, the system may automatically determine that the group should be generated in response to determining that the conditions are satisfied, provide a user interface output suggesting generation of the group, and then generate the group in response to a positive user interface received in response to the user interface output.
[0092] In block 254, the system obtains the processing capabilities of each assistant device in the group of assistant devices. For example, the system may be one of the assistant devices in the group. In this example, the assistant device may obtain its own processing capabilities, and the other assistant devices in the group may transfer their processing capabilities to the assistant device. As another example, the processing capabilities of the assistant device may be stored in the device topology, and the system may retrieve them from the device topology. As yet another example, the system may be a cloud-based component, and the assistant devices in the group may each transfer their processing capabilities to the system.
[0093] The processing capabilities of the assistant device may include a corresponding processor value based on the capabilities of one or more processors on the device, a corresponding memory value based on the size of the memory on the device, and / or a corresponding disk space value based on the available disk space. For example, the processor value may include details about one or more operating frequencies of the processor, details about the size of the processor's cache, whether each of the processors is a GPU, CPU, or DSP, and / or other details. As another example, the processor value may additionally or alternatively include a higher-level classification of the processor's capabilities, such as high, medium, or low, or GPU+CPU+DSP, high-power CPU+DSP, medium-power CPU+DSP, or low-power CPU+DSP. As another example, the memory value may include details of the memory (such as the specific size of the memory), or may include a higher-level classification of the memory, such as high, medium, or low. As yet another example, the disk space value may include details about the available disk space (such as the specific size of the disk space), or may include a higher-level classification of the available disk space, such as high, medium, or low.
[0094] In block 256, the system leverages the processing power of block 254 in determining a collective set of on-device models for the group. For example, the system can determine a collection of on-device models that seeks to maximize the use of collective processing power while ensuring that the on-device models in the collection can each be stored and used locally on the device that can store and use the on-device models. The system can also seek to ensure, if possible, that the selected collection includes a complete (or more complete than other candidate collections) pipeline of on-device models. For example, the system can select a collection that includes an ASR model but a less robust NLU model than a collection that includes a more robust NLU model but no ASR model.
[0095] In some embodiments, block 256 includes sub-block 256A, in which the system utilizes usage data when selecting an overall set of on-device models for the group. Past usage data can be data related to past assistant interactions at one or more assistant devices in the group and / or at one or more additional assistant devices in the ecosystem. In some embodiments, in sub-block 256A, the system considers usage data, along with the considerations mentioned above, when selecting on-device models to include in the set. For example, if processing capabilities allow for either a more accurate ASR model (over a less accurate ASR model) or a more robust NLU model (over a less robust NLU model) to be included in the set, but not both, usage data can be used to determine which to select. For example, if usage data indicates that past assistant interactions primarily (or exclusively) involved intents covered by a less robust NLU model (which can be included in the set with a more accurate ASR model), then the more accurate ASR model can be selected for inclusion in the set. On the other hand, if the usage data reflects that many of the included past assistant interactions involved intents covered by a more robust NLU model than a less robust NLU model, then the more robust NLU model can be selected for inclusion in the set. In some implementations, a candidate set is first determined based on processing power without considering usage data, and then, if multiple valid candidate sets exist, usage data can be used to select one model over others.
[0096] In block 258, the system causes each assistant device in the group to locally store a corresponding subset of the total set of on-device models. For example, the system can transmit corresponding instructions to each assistant device in the group regarding which on-device models should be downloaded. Based on the received instructions, each assistant device in the group can then download the corresponding model from a remote database. As another example, the system can retrieve the on-device model and push the corresponding on-device model to each assistant device in the group. As yet another example, any on-device model that was stored on one of the assistant devices in the group before adaptation and will be stored on another of the assistant devices in the group during adaptation can be directly transferred between the respective devices. For example, suppose a first assistant device stores an ASR model before adaptation, and during adaptation, the same ASR model is stored on a second assistant device and then cleared from the first assistant device. In this instance, the system can instruct the first assistant device to transfer the ASR model to the second assistant device for local storage there (and / or instruct the second assistant device to download it from the first assistant device), and the first assistant device can then clear the ASR model. In addition to avoiding WAN traffic, transmitting the pre-adaptation models locally can maintain any personalization of the on-device models that previously occurred at the time of transmission. The personalized models can be more accurate for users of the ecosystem compared to the non-personalized counterparts of the models in remote storage. As another example, for those assistant devices that included stored training instances before adaptation to personalize the on-device models on any assistant devices before adaptation, such training instances can be passed to assistant devices that will have corresponding models downloaded from a remote database after adaptation. The assistant devices that have the on-device models after adaptation can then utilize the training instances to personalize the corresponding models downloaded from the remote database. The corresponding models downloaded from the remote database can be different from (e.g., smaller or larger than) the counterparts on which the training instances were utilized before adaptation, but the training instances can still be used to personalize the different downloaded on-device models.
[0097] In block 260, the system assigns corresponding roles to each of the assistant devices. In some embodiments, assigning corresponding roles includes causing each of the assistant devices to download and / or implement an engine corresponding to an on-device model stored locally on the assistant device. The engine can each utilize the corresponding on-device model when performing corresponding processing roles (such as performing all or part of ASR, performing wake word recognition for at least some wake words, performing warm word recognition for certain warm words, and / or performing authentication). In some embodiments, one or more of the processing roles are executed only at the direction of the lead device in the group of assistant devices. For example, an NLU processing role performed by a given device using an on-device NLU model can be executed only in response to the lead device transmitting corresponding text for NLU processing and / or a specific command to cause NLU processing to occur to the given device. As another example, a warm word monitoring processing role performed by a given device using an on-device warm cue engine and an on-device warm cue model can be executed only in response to the lead device transmitting a command to cause warm word processing to occur to the given device. For example, the lead device can cause the given device to monitor for the spoken occurrence of the warm word "stop" in response to an alarm sounding on the lead device or another device in the group. In some embodiments, one or more processing roles can be executed, at least selectively, independently of any instructions from the guidance assistant device. For example, a warm thread monitoring role executed by a given device using an on-device warm thread engine and an on-device warm thread model can be executed continuously unless explicitly disabled by the user. As another example, a warm thread monitoring role executed by a given device using an on-device warm thread engine and an on-device warm thread model can be executed continuously or based on a monitoring condition detected locally at the given device.
[0098] In block 262, the system causes the spoken utterance detected at one or more of the devices in the group to be collaboratively processed locally at the assistant devices in the group according to their roles. Various non-limiting examples of such collaborative processing are described herein. For example, the example reference Figure 1B1 、 1B2 , 1B3, 1C and 1D descriptions.
[0099] In block 264, the system determines whether there are any changes to the group, such as adding a device to the group, removing a device from the group, or disabling the group. If not, the system continues to execute block 262. If yes, the system proceeds to block 266.
[0100] In box 266, the system determines whether the change to the group causes one or more assistant devices in the group to now be solo (i.e., no longer assigned to a group). If so, the system proceeds to box 268 and causes each of the solo devices to locally store the pre-group on-device model and assume its pre-group on-device processing role. In other words, if the device is no longer in the group, it can be caused to revert to a state prior to the adaptation performed in response to its inclusion in the group. By these and other means, after reverting to the state, the solo device can operate in a solo capacity to functionally process various assistant requests. Prior to reverting to the state, the solo device may not be able to functionally process any assistant requests or at least a smaller number of assistant requests than it could have processed before reverting to the state.
[0101] In block 270, the system determines whether two or more devices remain in the changed group. If so, the system returns to block 254 and performs another iteration of blocks 254, 256, 258, 260, and 262 based on the changed group. For example, if the changed group includes an additional assistant device without losing any previous assistant devices in the group, adaptation can be performed to account for the additional processing power of the additional assistant device. If the decision in block 270 is no, the group has been disbanded, and the system proceeds to block 272, where method 200 ends (until another group is generated).
[0102] Figure 3 is a flow chart illustrating an example method 300 that can be implemented by each of a plurality of assistant devices in a group when adapting an on-device model and / or processing role for an assistant device in the group. While the operations of method 300 are shown in a particular order, this is not intended to be limiting. One or more operations may be reordered, omitted, or added.
[0103] The operations of method 300 are a specific example of method 200 that can be performed by each of the assistant devices in the group. Therefore, the operations refer to the assistant device (such as Figure 1A 、 Figure 1B1 、 Figure 1B2 、 Figure 1B3 、 Figure 1C and Figure 1D Each of the assistant devices in the group can perform method 300 in response to receiving input indicating that it has been included in the group.
[0104] In block 352, the assistant device receives a grouping indication indicating that it has been included in the group. In block 352, the assistant device also receives identifiers of the other assistant devices in the group. For example, each identifier can be a MAC address, an IP address, a tag assigned to a device (e.g., assigned for use in a device topology), a serial number, or other identifier.
[0105] In optional block 354, the assistant device transmits data to the other assistant devices in the group. The data is transmitted to the other devices using the identifier received in block 352. In other words, the identifier can be a network address, or can be used to find a network address to which data is to be transmitted. The transmitted data can include one or more of the processing values described herein, another device identifier, and / or other data.
[0106] In optional block 356 , the assistant device receives the data transmitted by the other device in block 354 .
[0107] In block 358, based on the data optionally received in block 356 or the identifier received in block 352, the assistant device determines whether it is the dominant device. For example, the device may select itself as the dominant device if its own identifier is the lowest (or alternatively, the highest) value compared to the other identifiers received in block 352. As another example, the device may select itself as the dominant device if its processing value exceeds all other processing values received in the data in optional block 356. Other data may be transmitted in block 354 and received in block 356, and such other data may similarly enable an objective determination at the assistant device of whether it should be the dominant device. More generally, in block 358, the assistant device may utilize one or more objective criteria in determining whether it should be the dominant device.
[0108] In block 360, the assistant device determines whether it was determined to be the lead device in block 358. An assistant device that was not determined to be the lead device will then proceed to the “no” branch of block 360. An assistant device that was determined to be the lead device will proceed to the “yes” branch of block 360.
[0109] In the "yes" branch, in block 360, the assistant device utilizes the processing capabilities received from other assistant devices in the group as well as its own processing capabilities when determining the overall set of on-device models in the group. When optional block 354 is not executed or the data from block 354 does not include processing capabilities, processing capabilities can be transferred by other assistant devices to the lead device in optional block 352 or in block 370 (described below). In some embodiments, block 362 can be combined with Figure 2
[0066] Block 362 of method 200 may share one or more aspects in common. For example, in some embodiments, block 362 may also include considering past usage data when determining the overall set of on-device models.
[0110] In block 364, the assistant device transmits to each of the other assistant devices in the group a respective indication of the on-device models in the overall set that the other assistant devices are to download. In block 364, the assistant device may also optionally transmit to each of the other assistant devices in the group a respective indication of the processing role to be performed by the assistant device using the on-device models.
[0111] In block 366, the assistant device downloads and stores the on-device models in the collection that are assigned to the assistant device. In some implementations, blocks 364 and 366 may be associated with Figure 2 Block 258 of method 200 of FIG. 10 shares one or more aspects in common.
[0112] In block 368, the assistant device coordinates the collaborative processing of the assistant request, including utilizing its own on-device model in performing portions of the collaborative processing. In some implementations, block 368 may be associated with Figure 2 Block 262 of method 200 of FIG. 1 shares one or more aspects in common.
[0113] Turning now to the "no" branch, the assistant device transfers its processing capabilities to the lead device in optional block 370. For example, when block 354 is executed and the processing capabilities are included in the data transferred in block 354, block 370 may be omitted.
[0114] In block 372 , the assistant device receives from the lead device an indication of an on-device model to download and, optionally, an indication of a processing role.
[0115] In block 374, the assistant device downloads and stores the on-device model reflected in the indication of the on-device model received in block 372. In some implementations, blocks 372 and 374 can be combined with Figure 2 Block 258 of method 200 of FIG. 10 shares one or more aspects in common.
[0116] In block 376, the assistant device utilizes its on-device model when performing the assistant's portion of the collaborative processing request. In some implementations, block 376 may be associated with Figure 2 Block 262 of method 200 of FIG. 1 shares one or more aspects in common.
[0117] Figure 4is a block diagram of an example computing device 410 that can optionally be used to perform one or more aspects of the techniques described herein. In some implementations, one or more of the assistant device and / or other components can include one or more components of the example computing device 410.
[0118] Computing device 410 typically includes at least one processor 414 that communicates with a number of peripheral devices via a bus subsystem 412. These peripheral devices may include a storage subsystem 425 (including, for example, a memory subsystem 425 and a file storage subsystem 426), a user interface output device 420, a user interface input device 422, and a network interface subsystem 416. The input and output devices allow a user to interact with computing device 410. The network interface subsystem 416 provides an interface to an external network and is coupled to corresponding interface devices in other computing devices.
[0119] User interface input devices 422 may include a keyboard, a pointing device (such as a mouse, trackball, touchpad, or drawing tablet), a scanner, a touch screen incorporated into a display, an audio input device (such as a voice recognition system, a microphone), and / or other types of input devices. In general, the use of the term "input device" is intended to include all possible types of devices and ways of inputting information into computing device 410 or onto a communication network.
[0120] User interface output devices 420 may include a display subsystem, a printer, a fax machine, or a non-visual display (such as an audio output device). The display subsystem may include a cathode ray tube ("CRT"), a flat-panel device (such as a liquid crystal display ("LCD")), a projection device, or some other mechanism for creating a visible image. The display subsystem may also provide a non-visual display, such as via an audio output device. In general, the use of the term "output device" is intended to include all possible types of devices and ways of outputting information from computing device 410 to a user or to another machine or computing device.
[0121] The storage subsystem 425 stores programming and data structures that provide some or all of the functionality of the modules described herein. For example, the storage subsystem 425 may include logic for performing selected aspects of one or more of the methods described herein and / or implementing the various components depicted herein.
[0122] These software modules are typically executed by processor 414 alone or in combination with other processors. The memory used in storage subsystem 425 may include several memories, including main random access memory ("RAM") 430 for storing instructions and data during program execution and read-only memory ("ROM") 432 in which fixed instructions are stored. File storage subsystem 426 may provide persistent storage for program and data files and may include a hard drive, a floppy disk drive with associated removable media, a CD-ROM drive, an optical drive, or a removable media cartridge. Modules implementing the functionality of a particular embodiment may be stored by file storage subsystem 426 in storage subsystem 425 or in another machine accessible by processor 414.
[0123] The bus subsystem 412 provides a mechanism for the various components and subsystems of the computing device 410 to communicate with each other as intended. Although the bus subsystem 412 is shown schematically as a single bus, alternative implementations of the bus subsystem may use multiple busses.
[0124] The computing device 410 can be of different types, including a workstation, a server, a computing cluster, a blade server, a server farm, or any other data processing system or computing device. Due to the ever-changing nature of computers and networks, for the purposes of illustrating some embodiments, Figure 4 The description of computing device 410 depicted in FIG is intended only as a specific example. Figure 4 Many other configurations of computing device 410 are possible with more or fewer components than the computing device depicted in FIG.
[0125] Where the systems described herein collect or utilize personal information about users (or generally referred to herein as "participants"), users may be provided with the opportunity to control whether a program or feature collects user information (e.g., information about the user's social network, social actions or activities, occupation, preferences, or current geographic location) or to control whether and / or how content is received from content servers that may be more relevant to the user. Furthermore, before certain data is stored or used, the data may be processed in one or more ways to remove personally identifiable information. For example, the user's identity may be processed so that personally identifiable information about the user cannot be determined, or the user's geographic location (from which geographic location information (such as city, zip code, or state level) is derived) may be generalized so that the user's specific geographic location cannot be determined. Thus, users may have control over how information about the user is collected and / or used.
[0126] In some embodiments, a method is provided, including generating an assistant device group of different assistant devices. The different assistant devices include at least a first assistant device and a second assistant device. When generating the group, the first assistant device includes a first set of locally stored on-device models for use when locally processing assistant requests directed to the first assistant device. Further, when generating the group, the second assistant device includes a second set of locally stored on-device models for use when locally processing assistant requests directed to the second assistant device. The method also includes determining, based on corresponding processing capabilities of each of the different assistant devices in the assistant device group, a total set of locally stored on-device models for collaboratively locally processing assistant requests directed to any of the different assistant devices in the assistant device group. The method also includes, in response to generating the assistant device group, causing each of the different assistant devices to locally store a corresponding subset of the total set of locally stored on-device models, and assigning one or more corresponding processing roles to each of the different assistant devices in the assistant device group. Each of the processing roles utilizes one or more corresponding locally stored on-device models from the locally stored on-device models. Further, causing each assistant device in different assistant devices to locally store the corresponding subset includes: causing the first assistant device to clear one or more first-device models in the first set to provide storage space for the corresponding subset stored locally on the first assistant device, and causing the second assistant device to clear one or more second-device models in the second set to provide storage space for the corresponding subset stored locally on the second assistant device. The method also includes: after assigning the corresponding processing role to each assistant device in the different assistant devices in the assistant device group: detecting spoken utterances via a microphone of at least one of the different assistant devices in the assistant device group, and in response to the spoken utterances being detected via the microphone of the assistant device group, causing the spoken utterances to be collaboratively processed locally by the different assistant devices in the assistant device group using their corresponding processing roles.
[0127] These and other implementations of the technology disclosed herein can optionally include one or more of the following features.
[0128] In some embodiments, causing the first assistant device to clear one or more first-device models from the first set includes causing the first assistant device to clear a first-device wake-word detection model from the first set that is used to detect the first wake-word. In these embodiments, a corresponding subset stored locally on the second assistant device includes a second-device wake-word detection model that is used to detect the first wake-word, and assigning the corresponding processing role includes assigning the first wake-word detection role to the second assistant device, the first wake-word detection role utilizing the second-device wake-word detection model when monitoring for an occurrence of the first wake-word. In some of these embodiments, the spoken utterance includes the first wake-word followed by the assistant command, and in the first wake-word detection role, the second assistant device detects an occurrence of the first wake-word and, in response to detecting the occurrence of the first wake-word, causes execution of an additional processing role from the corresponding processing role. In some versions of these embodiments, the additional processing role from the corresponding processing role is executed by the first assistant device, and the second assistant device causes execution of the additional processing role from the corresponding processing role by transmitting an indication of detection of the first wake-word to the first assistant device.
[0129] In some embodiments, the corresponding subset stored locally on the first assistant device includes a first device first wake-up word detection model used to detect one or more first wake-up words, and does not include any wake-up word detection model used to detect one or more second wake-up words. In some of these embodiments, the corresponding subset stored locally on the second assistant device includes a second device second wake-up word detection model used to detect one or more second hot words, and does not include any wake-up word detection model used to detect one or more first wake-up words. In some versions of these embodiments, assigning the corresponding processing role includes: assigning a first wake-up word detection role to the first assistant device, the first wake-up word detection role utilizing the first device wake-up word detection model when monitoring for the occurrence of one or more first wake-up words; and assigning a second wake-up word detection role to the second assistant device, the second wake-up word detection role utilizing the second device wake-up word detection model when monitoring for the occurrence of one or more second wake-up words.
[0130] In some embodiments, the corresponding subset stored locally on the first assistant device includes a first language speech recognition model used to perform speech recognition in the first language and does not include any speech recognition model used to recognize speech in the second language. In some of these embodiments, the corresponding subset stored locally on the second assistant device includes a second language speech recognition model used to perform speech recognition in the second language and does not include any speech recognition model used to recognize speech in the second language. In some versions of these embodiments, assigning the corresponding processing roles includes: assigning a first language speech recognition role to the first assistant device, the first language speech recognition role utilizing the first language speech recognition model when performing speech recognition in the first language; and assigning a second language speech recognition role to the second assistant device, the second language speech recognition role utilizing the second language speech recognition model when performing speech recognition in the second language.
[0131] In some embodiments, the corresponding subset stored locally on the first assistant device includes the first portion of the speech recognition model used to perform the first portion of speech recognition and does not include the second portion of the speech recognition model. In some of these embodiments, the corresponding subset stored locally on the second assistant device includes the second portion of the speech recognition model used to perform the second portion of speech recognition and does not include the first portion of the speech recognition model. In some versions of these embodiments, assigning the corresponding processing roles includes: assigning a first portion of a language speech recognition role to the first assistant device, the first portion of the language speech recognition role utilizing the first portion of the speech recognition model when generating a corresponding embedding for the corresponding speech and transmitting the corresponding embedding to the second assistant device; and assigning a second language speech recognition role to the second assistant device, the second language speech recognition role utilizing the corresponding embedding and the second language speech recognition model from the first assistant device when generating a corresponding recognition of the corresponding speech.
[0132] In some embodiments, the corresponding subset stored locally on the first assistant device includes a speech recognition model used to perform a first portion of speech recognition. In some of these embodiments, assigning the corresponding processing role includes: assigning a first portion of a language speech recognition role to the first assistant device, the first portion of the language speech recognition role utilizing the speech recognition model when generating an output and transmitting the corresponding output to the second assistant device; and assigning a second language speech recognition role to the second assistant device, the second language speech recognition role performing a beam search on the corresponding output from the first assistant device when generating a corresponding recognition of the corresponding speech.
[0133] In some embodiments, the corresponding subset stored locally on the first assistant device includes one or more pre-adapted natural language understanding models used to perform semantic analysis of natural language input, and the one or more initial natural language understanding models occupy a first amount of local disk space at the first assistant device. In some of these embodiments, the corresponding subset stored locally on the first assistant device includes one or more adapted natural language understanding models, the one or more adapted natural language understanding models including at least one additional natural language understanding model in addition to the one or more initial natural language understanding models, and occupies a second amount of local disk space at the first assistant device, the second amount being greater than the first amount.
[0134] In some embodiments, the corresponding subset stored locally on the first assistant device includes the first device natural language understanding model used for semantic analysis for one or more first categories, and does not include any natural language understanding model used for semantic analysis for the second category. In some of these embodiments, the corresponding subset stored locally on the second assistant device includes the second device natural language understanding model used for semantic analysis for at least the second category.
[0135] In some implementations, the corresponding processing capabilities of each assistant device in the different assistant devices in the assistant device group include a corresponding processor value based on the capabilities of the processor on one or more devices, a corresponding memory value based on the size of the memory on the device, and a corresponding disk space value based on the available disk space.
[0136] In some implementations, generating the assistant device group of different assistant devices is in response to a user interface input that explicitly indicates a desire to group the different assistant devices.
[0137] In some implementations, generating the assistant device group of different assistant devices is performed automatically in response to determining that the different assistant devices satisfy one or more proximity conditions relative to each other.
[0138] In some implementations, generating an assistant device group of different assistant devices is performed in response to affirmative user interface input received in response to a suggestion to create an assistant device group, and the suggestion is automatically generated in response to determining that the different assistant devices satisfy one or more proximity conditions relative to each other.
[0139] In some embodiments, the method further includes: after assigning the corresponding processing role to each of the different assistant devices in the assistant device group, determining that the first assistant device is no longer in the group; and in response to determining that the first assistant device is no longer in the group, causing the first assistant device to replace the corresponding subset stored locally on the first assistant device with the first on-device model in the first set.
[0140] In some implementations, determining the overall set is further based on usage data reflecting past usage at one or more of the assistant devices in the group.
[0141] In some of these embodiments, determining the overall set includes: determining multiple candidate sets based on the corresponding processing capabilities of each assistant device in different assistant devices in the assistant device group, and the multiple candidate sets can be collectively stored and collectively used locally by the assistant devices in the group; and selecting the overall set from the candidate sets based on the usage data.
[0142] In some embodiments, a method implemented by one or more processors of an assistant device is provided. The method includes: in response to determining that the assistant device is included in an assistant device group, the assistant device group including the assistant device and one or more additional assistant devices, determining that the assistant device is a lead device for the group. The method also includes: in response to determining that the assistant device is the lead device in the group, based on the processing capabilities of the assistant device and based on the received processing capabilities of each of the one or more additional assistant devices: determining a total set of on-device models for collaboratively locally processing assistant requests directed to any of the different assistant devices in the assistant device group; and, for each of the on-device models, determining a corresponding designation of which of the assistant devices in the group will locally store the on-device model. The method also includes: in response to determining that the assistant device is the lead device in the group: communicating with the one or more additional assistant devices to cause the one or more additional assistant devices to respectively locally store any of the on-device models with corresponding designations for the additional assistant devices; locally storing the on-device models with corresponding designations for the assistant devices at the assistant device; and assigning one or more corresponding processing roles to each of the assistant devices in the group for collaborative local processing of assistant requests directed to the group.
[0143] These and other implementations of the technology disclosed herein can optionally include one or more of the following features.
[0144] In some implementations, determining that the assistant device is the lead device in the group includes: comparing the processing capability of the assistant device to the received processing capability of each of the one or more additional assistant devices; and determining that the assistant device is the lead device in the group based on the comparison.
[0145] In some implementations, an assistant device group is created in response to a user interface input that explicitly indicates a desire to group different assistant devices.
[0146] In some embodiments, the method further includes, in response to determining that the assistant device is the dominant device in the group and in response to receiving an assistant request at one or more of the assistant devices in the group, coordinating collaborative local processing of the assistant request using corresponding processing roles assigned to the assistant devices.
[0147] In some embodiments, a method implemented by one or more processors of an assistant device is provided and includes determining that the assistant device has been removed from a group of different assistant devices. The group is a group that already includes the assistant device and at least one additional assistant device. When the assistant device is removed from the group, the assistant device locally stores a collection of on-device models, and the collection of on-device models is insufficient to fully process spoken utterances directed to the automated assistant locally at the assistant device. The method also includes: in response to determining that the assistant device has been removed from the group of assistant devices: causing the assistant device to clear one or more of the on-device models in the collection and to retrieve and locally store one or more additional on-device models. After retrieving and locally storing the one or more additional on-device models of the assistant device, the one or more additional on-device models and any remaining on-device models in the collection can be used to fully process spoken utterances directed to the automated assistant locally at the assistant device.
Claims
1. A method implemented by one or more processors of an assistant device, the method comprising: In response to determining that the assistant device is included in an assistant device group, the assistant device group including the assistant device and one or more additional assistant devices: determining that the assistant device is a lead device for the group; In response to determining that the assistant device is the lead device of the group: Based on the processing capability of the assistant device and based on the received processing capability of each of the one or more additional assistant devices: determining a total set of on-device models for collaboratively and locally processing an assistant request directed to any one of different assistant devices in the group of assistant devices; as well as For each of the on-device models, determining a corresponding designation of which assistant device in the group of assistant devices will locally store the on-device model; communicating with the one or more additional assistant devices so that the one or more additional assistant devices each locally stores any one of the on-device models having a corresponding designation for the additional assistant device; storing locally at the assistant device an on-device model with corresponding designations for the assistant device; as well as Each of the assistant devices in the group is assigned one or more corresponding processing roles for collaborative local processing of assistant requests directed to the group.
2. The method according to claim 1, wherein Determining that the assistant device is the lead device of the group includes: comparing the processing capability of the assistant device to the received processing capability of each of the one or more additional assistant devices; and The assistant device is determined to be the lead device of the group based on the comparison.
3. The method according to claim 2, wherein: The processing capabilities of each of the one or more additional assistant devices are each transmitted to the assistant device by a corresponding one of the one or more additional assistant devices.
4. The method according to claim 1, wherein The assistant device group is created in response to a user interface input that explicitly indicates a desire to group assistant devices.
5. The method according to claim 1, further comprising: In response to determining that the assistant device is the lead device of the group, and in response to receiving an assistant request at one or more assistant devices of the group: Coordinated local processing of the assistant request is coordinated using corresponding processing roles assigned to the assistant devices.
6. The method according to claim 1, wherein The assistant device group is automatically created in response to determining that the assistant device and the one or more additional assistant devices satisfy one or more proximity conditions relative to each other.
7. The method according to claim 6, wherein: The one or more proximity conditions include: the assistant device and the one or more additional assistant devices being assigned to a same area within a given structure in a device topology.
8. The method according to claim 6, wherein: The one or more proximity conditions include: the assistant device and the one or more additional assistant devices all detecting an occurrence of a wake word at or near the same time.
9. The method according to claim 6, wherein: The one or more proximity conditions include: the assistant device transmitting a signal and the one or more additional assistant devices all detecting the presence of the signal when the signal is transmitted.
10. The method according to claim 1, wherein Assigning the one or more corresponding processing roles to each of the assistant devices in the group for collaborative local processing of assistant requests directed to the group, including: assigning an automatic speech recognition role to the assistant device, and A natural language understanding role is assigned to a given one of the one or more additional assistant devices.
11. An assistant device, comprising: microphone; speaker; a memory for storing instructions; one or more processors operable to execute the instructions to: In response to determining that the assistant device is included in an assistant device group, the assistant device group including the assistant device and one or more additional assistant devices: determining that the assistant device is a lead device for the group; In response to determining that the assistant device is the lead device of the group: Based on the processing capability of the assistant device and based on the received processing capability of each of the one or more additional assistant devices: determining a total set of on-device models for collaboratively and locally processing an assistant request directed to any one of different assistant devices in the group of assistant devices; as well as For each of the on-device models, determining a corresponding designation of which assistant device in the group of assistant devices will locally store the on-device model; communicating with the one or more additional assistant devices so that the one or more additional assistant devices each locally stores any one of the on-device models having a corresponding designation for the additional assistant device; storing locally at the assistant device an on-device model with corresponding designations for the assistant device; as well as Each of the assistant devices in the group is assigned one or more corresponding processing roles for collaborative local processing of assistant requests directed to the group.
12. The assistant device according to claim 11, wherein: Upon determining that the assistant device is the lead device of the group, one or more of the processors are configured to: comparing the processing capability of the assistant device to the received processing capability of each of the one or more additional assistant devices; as well as The assistant device is determined to be the lead device of the group based on the comparison.
13. The assistant device according to claim 12, wherein: The processing capabilities of each of the one or more additional assistant devices are each transmitted to the assistant device by a corresponding one of the one or more additional assistant devices.
14. The assistant device according to claim 11, wherein: The assistant device group is created in response to a user interface input that explicitly indicates a desire to group assistant devices.
15. The assistant device according to claim 11, wherein: One or more of the processors are further operable to execute instructions for: In response to determining that the assistant device is the lead device of the group, and in response to receiving an assistant request at one or more assistant devices of the group: Coordinated local processing of the assistant request is coordinated using corresponding processing roles assigned to the assistant devices.
16. The assistant device according to claim 11, wherein: The assistant device group is automatically created in response to determining that the assistant device and the one or more additional assistant devices satisfy one or more proximity conditions relative to each other.
17. The assistant device according to claim 16, wherein: The one or more proximity conditions include: the assistant device and the one or more additional assistant devices being assigned to a same room within a given structure in a device topology.
18. The assistant device according to claim 16, wherein: The one or more proximity conditions include: the assistant device and the one or more additional assistant devices all detecting an occurrence of a wake word at or near the same time.
19. The assistant device according to claim 16, wherein: The one or more proximity conditions include: the assistant device transmitting a signal and the one or more additional assistant devices all detecting the presence of the signal when the signal is transmitted.
20. The assistant device according to claim 11, wherein: Assigning the one or more corresponding processing roles to each of the assistant devices in the group for collaborative local processing of assistant requests directed to the group, one or more of the processors being configured to: A natural language understanding role is assigned to a given one of the one or more additional assistant devices.
21. A method implemented by one or more processors of an assistant device, the method comprising: Determining that a first assistant device has been removed from a group of different assistant devices, the group of different assistant devices already including the first assistant device and a second assistant device, wherein when the first assistant device is removed from the group of different assistant devices: the first assistant device locally storing a first on-device model set, the first on-device model set being insufficient to fully process spoken utterances directed to the automated assistant locally at the first assistant device; and In response to determining that the first assistant device has been removed from the set of different assistant devices: Cause the first assistant device to clear one or more on-device models from the on-device models in the first on-device model set and retrieve and locally store one or more additional on-device models, wherein, after the one or more additional on-device models are retrieved and locally stored at the first assistant device, the one or more additional on-device models and any remaining on-device models in the first on-device model set can be used to fully process the spoken utterance directed to the automated assistant locally at the first assistant device.
22. The method according to claim 21, wherein The set of different assistant devices further includes a third assistant device.
23. The method of claim 22, further comprising: determining a total set of on-device models, the total set of on-device models comprising one or more on-device models from the first on-device model set, one or more on-device models from the second on-device model set, and one or more on-device models from the third on-device model set, wherein, when the first assistant device has been removed from the set of different assistant devices, the model set on the second device is stored locally at the second assistant device, and Wherein, when the first assistant device has been removed from the group of different assistant devices, the model set on the third device is locally stored at the third assistant device.
24. The method of claim 23, further comprising: causing the second assistant device to locally store a first subset of the overall set of on-device models, wherein the first subset of the overall set of on-device models is different from the second on-device set of models; and The third assistant device is caused to locally store a second subset of the overall set of on-device models, wherein the second subset of the overall set of on-device models is different from the third set of on-device models.
25. The method according to claim 24, wherein A particular on-device model in the first subset of the overall set of on-device models has a processing capability that corresponds to a processing capability of one or more of the on-device models in the first set of on-device models.
26. The method of claim 24, further comprising: In response to causing the second assistant device to locally store the first subset of the overall set of on-device models, and causing the third assistant device to locally store the second subset of the overall set of on-device models: assigning a first processing role to the second assistant device, wherein the first processing role is based on the first subset of the total set of on-device models stored locally at the second assistant device; and A second processing role is assigned to the third assistant device, wherein the second processing role is based on the second subset of the total set of on-device models stored locally at the third assistant device.
27. The method of claim 26, further comprising: causing the additional spoken utterance to be collaboratively processed by the set of different assistant devices, wherein causing the additional spoken utterance to be collaboratively processed by the set of different assistant devices comprises: causing the second assistant device to perform the first processing role; and The third assistant device is caused to perform the second processing role.
28. A system comprising: a memory for storing instructions; as well as One or more processors operable to execute the instructions for: Determining that a first assistant device has been removed from a group of different assistant devices, the group of different assistant devices already including the first assistant device and a second assistant device, wherein when the first assistant device is removed from the group of different assistant devices: the first assistant device locally storing a first on-device model set, the first on-device model set being insufficient to fully process spoken utterances directed to the automated assistant locally at the first assistant device; and In response to determining that the first assistant device has been removed from the set of different assistant devices: Cause the first assistant device to clear one or more on-device models from the on-device models in the first on-device model set and retrieve and locally store one or more additional on-device models, wherein, after the one or more additional on-device models are retrieved and locally stored at the first assistant device, the one or more additional on-device models and any remaining on-device models in the first on-device model set can be used to fully process the spoken utterance directed to the automated assistant locally at the first assistant device.