Dynamically adapting the on-device model of a grouping assistant device for collaborative processing of assistant requests
By dynamically adapting the models and processing roles on the assistant device, the problem of limited processing capabilities of the assistant device in the prior art is solved, and higher robustness and accuracy are achieved, reducing latency and network usage.
Patent Information
- Application Number
- CN202080100916.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-11-13
- Filing Date
- 2020-12-11
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2040-12-11
AI Technical Summary
Existing automation assistant devices, due to limited processing capabilities, result in locally implemented components that are unreliable and accurate, especially on older or less costly devices.
By dynamically adapting models and processing roles on the assistant device, the models and processing roles on the device are determined based on the individual processing capabilities in the group, so as to improve the collective robustness and ability of models and processing roles.
Improves the robustness and accuracy of the pipelines on the equipment of the assistant component, reduces latency, improves the security of user data, reduces network usage, and reduces the amount of data transmission of the remote automation assistant component.
Smart Images

Figure CN115668124B_ABST
Abstract
Description
Background Art
[0001] A person can participate in a human-machine conversation with an interactive software application referred to herein as an "automation assistant" (also known as a "chatbot", "interactive personal assistant", "intelligent personal assistant", "personal voice assistant", "conversational agent", etc.). For example, a person (who can be referred to as a "user" when interacting with the automation assistant) can use spoken natural language input (i.e., spoken utterances) to provide commands and / or requests to the automation assistant, which in some cases can be converted to text and then processed. The commands and / or requests can additionally or alternatively be provided via one or more other input modalities, such as text (e.g., typed) natural language input, touchscreen input, and / or touchless gesture input (e.g., detected by a camera of the corresponding assistant device). The automation assistant typically responds to commands or requests by providing responsive user interface outputs (e.g., audible and / or visual user interface outputs), controlling smart devices, and / or performing other actions.
[0002] Automation assistants typically rely on a pipeline of components when processing user requests. For example, a wake word detection engine can be used to process audio data while monitoring for the occurrence of a spoken wake word (e.g., "OK Assistant"), and in response to detecting the occurrence, cause the processing of other components to occur. As another example, an automatic speech recognition (ASR) engine can be used to process audio data including spoken utterances to generate a transcription of the user's utterance (i.e., a sequence of terms and / or other tokens). The ASR engine can process the audio data based on the subsequent occurrence of the spoken wake word detected by the wake word detection engine and / or in response to other invocations of the automation assistant. As another example, a natural language understanding (NLU) engine can be used to process the text of the request (e.g., text converted from a spoken utterance using ASR) to generate a symbolic representation or belief state that is a semantic representation of the text. For example, the belief state can include the intent corresponding to the text and optionally include parameters of that intent (e.g., slot values). Once fully formed through one or more conversation turns (e.g., all mandatory parameters have been parsed), the belief state representation represents the action to be performed in response to the spoken utterance. Then, a separate fulfillment component can utilize the fully formed belief state to perform the action corresponding to the belief state.
[0003] The user utilizes one or more assistant devices (client devices having an automation assistant interface) when interacting with the automation assistant. The pipeline of components for processing requests provided at the assistant device can include components executed locally at the assistant device and / or components implemented at one or more remote servers that communicate with the assistant device over a network.
[0004] Efforts have been made to increase the number of components executed locally at the assistant device and / or improve the robustness and / or accuracy of such components. The motivation for these efforts is consideration, such as reducing latency, improving data security, reducing network usage, and / or seeking to achieve other technical benefits. As an example, some assistant devices may include a local wake word engine and / or a local ASR engine.
[0005] However, due to the limited processing capabilities of various assistant devices, components implemented locally at the assistant device may be less robust and / or accurate compared to their cloud-based counterparts. This may be particularly true for older and / or less expensive assistant devices, which may lack: (a) the processing power and / or memory capacity to execute various components and / or utilize their associated models; (b) and / or the disk space capacity to store various associated models. Summary of the Invention
[0006] Embodiments disclosed herein relate to dynamically adapting which on-device models are stored locally at an assistant device in a group of assistant devices and / or adapting the assistant processing roles of the assistant devices in the group of assistant devices. In some of these embodiments, for each assistant device in the group of assistant devices, the corresponding on-device model and / or the corresponding processing role are determined based on collectively considering the individual processing capabilities of the assistant devices in the group. For example, based on the individual processing capabilities of a given assistant device (e.g., whether the given assistant device can store those on-device models and execute those processing roles given processor, memory, and / or storage constraints) and in view of the corresponding processing capabilities of the other assistant devices in the group (e.g., whether the other devices are capable of storing other necessary on-device models and / or executing other necessary processing roles), the on-device model and / or the processing role of the given assistant device can be determined. In some embodiments, usage data may also be used to determine the corresponding on-device model and / or the corresponding processing role for each assistant device in the group.
[0007] Embodiments disclosed herein additionally or alternatively relate to collaboratively utilizing a group of assistant devices and their associated post-adaptation on-device models and / or post-adaptation processing roles when collaboratively processing an assistant request directed to any one of the assistant devices in the group.
[0008] In these and other ways, given the processing capabilities of these assistant devices, the on-device models and on-device processing roles can be distributed among multiple different assistant devices in a group. Further, when distributed among multiple assistant devices in a group, the collective robustness and / or capabilities of the on-device models and / or on-device processing roles exceed the robustness and / or capabilities that any one of the assistant devices might have individually. In other words, the embodiments disclosed herein can effectively implement an on-device pipeline of assistant components distributed among the assistant devices in a group. The robustness and / or accuracy of such a distributed pipeline far exceed the robustness and / or accuracy capabilities of any pipeline implemented alternatively on only a single assistant device in the group of assistant devices. Improving the robustness and / or accuracy disclosed herein can reduce latency for a greater number of assistant requests. Further, the improved robustness and / or accuracy results in less (or even no) data being transmitted to a remote automated assistant component to parse an assistant request. This directly leads to increased security of user data, reduced frequency of network usage, reduced amount of data transmitted over the network, and / or reduced latency in parsing an assistant request (e.g., these assistant requests parsed locally can be parsed with less latency compared to cases involving a remote assistant component).
[0009] Certain adaptations may cause one or more assistant devices in the group to lack the engines and / or models necessary for that assistant device to process many assistant requests on its own. For example, certain adaptations may cause an assistant device to lack any wake word detection capabilities and / or ASR capabilities. However, when in a group and adapted according to the embodiments disclosed herein, the assistant device can cooperate with other assistant devices in the group to collaboratively process assistant requests, with each other assistant device performing its own processing role and leveraging its own on-device model when doing so. Thus, assistant requests directed to that assistant device can still be processed collaboratively with other assistant devices in the group.
[0010] In various embodiments, the adaptation of the assistant devices in a group is performed in response to generating the group or modifying the group (e.g., incorporating or removing an assistant device from the group). As described herein, a group can be generated based on explicit user input and / or automatically generated based on, for example, determining that the assistant devices in the group meet a proximity condition relative to each other, the explicit user input indicating the desire for the group. In embodiments where the adaptation is performed only when the group is created in this manner, the occurrence of two assistant requests received simultaneously at two separate devices in the group (and potentially not capable of being collaboratively processed in parallel) can be mitigated. For example, when the proximity condition is considered during group generation, two unrelated simultaneous requests are less likely to be received at two different assistant devices in the group. For example, this occurrence is less likely when the assistant devices in the group are all in the same room compared to when the assistant devices in the group are spread across multiple floors of a house. As another example, when the user input explicitly indicates that a group should be generated, non-overlapping assistant requests are likely to be provided to the assistant devices in the group.
[0011] As a specific example in various embodiments, assume that a group of assistant devices including a first assistant device and a second assistant device is generated. Further assume that, when the group of assistant devices is generated, the first assistant device includes a wake word engine and a corresponding wake word model, a warm cue engine and a corresponding warm cue model, an authentication engine and a corresponding authentication model, and a corresponding on-device ASR engine and a corresponding ASR model. Further assume that, when the group of assistant devices is generated, the second assistant device also includes the same engines and models (or variants thereof) as the first assistant device, and additionally includes: an on-device NLU engine and a corresponding NLU model, an on-device fulfillment and a corresponding fulfillment model, and an on-device TTS engine and a corresponding TTS model.
[0012] In response to the first assistant device and the second assistant device being grouped, the on-device models stored locally at the first assistant device and the second assistant device and / or the corresponding processing roles of the first assistant device and the second assistant device can be adapted. For each of the first assistant device and the second assistant device, the on-device model and the processing role can be determined based on considering the first processing capabilities of the first assistant device and the second assistant device. For example, a set of on-device models can be determined, and the set of on-device models includes a first subset that can be stored and utilized by the corresponding processing role / engine on the first assistant device. Further, the set can include a second subset that can be stored and utilized by the corresponding processing role on the second assistant device. For example, the first subset may only include an ASR model, but the ASR model of the first subset may be more robust and / or accurate than the pre-adaptation ASR model of the first assistant device. Further, they may need to utilize more computing resources when performing ASR. However, the first assistant device may only have the ASR model and the ASR engine of the first subset, and the pre-adaptation model and engine can be cleared, thereby releasing computing resources for utilization when performing ASR using the ASR model of the first subset. Continuing with the example, the second subset can include the same models as the previously included second assistant device, except that the ASR model can be omitted and a more robust and / or more accurate NLU model can replace the pre-adaptation NLU model. The more robust and / or more accurate NLU model may require more resources than the pre-adaptation NLU model, but these resources can be released by clearing the pre-adaptation ASR model (and omitting any ASR models from the second subset).
[0013] Then, in collaboratively processing an assistant request directed to any one of the assistant devices in the group, the first assistant device and the second assistant device can collaboratively utilize their associated post-adaptation on-device models and / or post-adaptation processing roles. For example, assume the spoken utterance "OK Assistant, increase the temperature two degrees". The wake-up cue engine of the second assistant device can detect the occurrence of the wake-up cue "OK Assistant". In response, the wake-up cue engine can transmit a command to the first assistant device to cause the ASR engine of the first assistant device to perform speech recognition on the audio data captured after the wake-up cue. The transcription generated by the ASR engine of the first assistant device can be transmitted to the second assistant device, and the NLU engine of the second assistant device can perform NLU on the transcription. The result of the NLU can be passed to the fulfillment engine of the second assistant device, and the fulfillment engine can use these NLU results to determine the command to be transmitted to the smart thermostat to cause it to increase the temperature by two degrees.
[0014] The foregoing is provided as an overview of only some embodiments. These and other embodiments are disclosed in more detail herein.
[0015] Additionally, some embodiments may include a system that includes one or more user devices, each user device having one or more processors and a memory operably coupled to the one or more processors, wherein the memory of the one or more user devices stores instructions that, responsive to execution of the instructions by the one or more processors of the one or more user devices, cause the one or more processors to perform any of the methods described herein. Some embodiments also include at least one non-transitory computer-readable medium that includes instructions that, responsive to execution of the instructions by one or more processors, cause the one or more processors to perform any of the methods described herein.
[0016] It should be understood that all combinations of the foregoing concepts and additional concepts described in more detail herein are contemplated as part of the subject matter disclosed herein. For example, all combinations of the appended claims of this disclosure are contemplated as part of the subject matter disclosed herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1A is a block diagram of an example assistant ecosystem according to embodiments disclosed herein in which no assistant devices are grouped and adapted.
[0018] Figure 1B1 、 1B2 and 1B3 respectively illustrate Figure 1A example assistant ecosystems in which a first assistant device and a second assistant device have been grouped and have different examples of adaptations that may be implemented.
[0019] Figure 1C illustrates Figure 1A example assistant ecosystems in which a first assistant device, a second assistant device, and a third assistant device have been grouped and have examples of adaptations that may be implemented.
[0020] Figure 1D illustrates Figure 1A example assistant ecosystems in which a third assistant device and a fourth assistant device have been grouped and have examples of adaptations that may be implemented.
[0021] Figure 2 is a flowchart of an example method illustrating an on-device model and / or processing role of assistant devices in an adaptation group.
[0022] Figure 3 is a flowchart of an example method that may be implemented by each of a plurality of assistant devices in a group when illustrating an on-device model and / or processing role of assistant devices in an adaptation group.
[0023] Figure 4 Illustrates an example architecture of a computing device. Detailed implementation
[0024] Many users can access the automated assistant using any one of multiple assistant devices. For example, some users can handle a coordinated "ecosystem" of assistant devices that can receive user input directed to the automated assistant and / or can be controlled by the automated assistant, such as one or more smart phones, one or more tablet computers, one or more vehicle computing systems, one or more wearable computing devices, one or more smart TVs, one or more interactive standalone speakers, one or more interactive standalone speakers with a display, one or more IoT devices, and other assistant devices.
[0025] Users can use any one of these assistant devices to participate in a human-machine conversation with the automated assistant (assuming the automated assistant client is installed and the assistant device is capable of receiving input). In some cases, these assistant devices can be scattered around the user's primary residence, secondary residence, workplace, and / or other structures. For example, mobile assistant devices (such as smart phones, tablet computers, smart watches, etc.) can be on the user's person and / or where the user last placed them. Other assistant devices (such as traditional desktop computers, smart TVs, interactive standalone speakers, and IoT devices) can be more stationary, yet may be located at various locations (e.g., rooms) within the user's home or workplace.
[0026] Initially turn to Figure 1A, an example assistant ecosystem is illustrated. The example assistant ecosystem includes a first assistant device 110A, a second assistant device 110B, a third assistant device 110C, and a fourth assistant device 110D. The assistant devices 110A to 110D can all be set within a home, enterprise, or other environment. Further, the assistant devices 110A to 110D can all be linked together in one or more data structures, or otherwise associated with each other. For example, the four assistant devices 110A to 110D can all be registered to the same user account, to the same set of user accounts, to a particular structure, and / or all be assigned to a particular structure in a device topology representation. For each of the assistant devices 110A to 110D, the device topology representation can include a corresponding unique identifier and can optionally include corresponding unique identifiers of other devices that are not assistant devices (but can interact via the assistant devices), such as IoT devices that do not include an assistant interface. Further, the device topology representation can specify device attributes associated with the respective assistant devices 110A to 110D. The device attributes of a given assistant device can indicate, for example, one or more input and / or output modalities supported by the respective assistant device, the processing capabilities of the respective assistant device, the brand, model, and / or unique identifier (such as a serial number) of the respective assistant device (based on which the processing capabilities can be determined), and / or other attributes. As another example, the four assistant devices can all be linked together, or otherwise associated with each other, according to being connected to the same wireless network (such as a secure access wireless network) and / or according to collectively communicating peer-to-peer with each other (e.g., via Bluetooth and after pairing). In other words, in some embodiments, multiple assistant devices can be considered linked together according to making secure network connections with each other and without being associated with each other in any data structure, and potentially adapted according to the embodiments disclosed herein.
[0027] As a non-limiting working example, the first assistant device 110A can be a first type of assistant device, such as a specific model of an interactive standalone speaker having a display and a camera. The second assistant device 110B can be a second type of assistant device, such as a first model of an interactive standalone speaker without a display or a camera. The assistant devices 110C and 110D can be a third type of assistant device, such as a third model of an interactive standalone speaker without a display. The third type (assistant devices 110C and 110D) can have less processing power compared to the second type (assistant device 110D). For example, the processor of the third type can have less processing power compared to the processor of the second type. For example, the processor of the third type may lack any GPU, while the processor of the first type includes a GPU. Also, for example, the processor of the third type can have a smaller cache and / or a lower operating frequency compared to the processor of the second type. As another example, the size of the on-device memory of the third type of device can be smaller than the size of the on-device memory of the second device (e.g., 1GB compared to 2GB). As yet another example, the available disk space of the third type can be smaller than the available disk space of the first type. The available disk space can be different from the currently available disk space. For example, the available disk space can be determined as the currently available disk space plus the disk space currently occupied by one or more on-device models. As another example, the available disk space can be the total disk space minus any space occupied by the operating system and / or other specific software. Continuing with this working example, the first type and the second type can have the same processing power.
[0028] In addition to being linked together in a data structure, two or more (e.g., all) of the assistant devices 110A to 110D also communicate with each other partially selectively via a local area network (LAN) 108. The LAN 108 can include a wireless network such as using Wi-Fi, a direct peer-to-peer network such as using Bluetooth, and / or other communication topologies using other communication protocols.
[0029] The assistant device 110A includes an assistant client 120A, which can be a standalone application above the operating system or can form all or part of the operating system of the assistant device 110A. In Figure 1AAmong them, the assistant client 120A includes a wake-up / call engine 121A1 and one or more associated on-device wake-up / call models 131A1. The wake-up / call engine 121A1 can monitor the occurrence of one or more wake-up or call cues, and in response to detecting one or more of the cues, can call one or more previously inactive functions of the assistant client 120A. For example, calling the assistant client 120A can include activating the ASR engine 122A1, the NLU engine 123A1, and / or other engines. For example, it can cause the ASR engine 122A1 to process other audio data frames after the wake-up or call cue (whereas before the call, no other processing of audio data frames occurred) and / or can cause the assistant client 120A to transmit other audio data frames and / or other data to be transmitted to the cloud-based assistant component 140 for processing (e.g., the audio data frames are processed by the remote ASR engine of the cloud-based assistant component 140).
[0030] In some embodiments, the wake-up cue engine 121A can continuously process (e.g., if not in an "inactive" mode) a stream of audio data frames based on the output from one or more microphones of the client device 110A to monitor for the occurrence of an oral wake word or call phrase (e.g., "OK Assistant", "Hey Assistant"). The processing can be performed by the wake-up cue engine 121A using one or more of the wake-up models 131A1. For example, one of the wake-up models 131A1 can be a neural network model trained to process audio data frames and generate an output indicating whether one or more wake words are present in the audio. When monitoring for the occurrence of a wake word, the wake-up cue engine 121 discards (e.g., after temporarily storing in a buffer) any audio data frames that do not include the wake word. As a supplement or alternative to monitoring for the occurrence of a wake word, the wake-up cue engine 121A1 can monitor for the occurrence of other call cues. For example, the wake-up cue engine 121A1 can also monitor the pressing of a call hardware button and / or a call software button. As another example, and continuing with the working example, when the assistant device 110A includes a camera, the wake-up cue engine 121A1 can also optionally process image frames from the camera to monitor for the occurrence of a call gesture (such as a wave while the user's gaze is directed at the camera) and / or other call cues (such as the user's gaze directed at the camera along with an indication that the user is speaking).
[0031] In Figure 1AIn [the context], the assistant client 120A further includes an automatic speech recognition (ASR) engine 122A1 and one or more associated on-device ASR models 132A1. The ASR engine 122A1 can be used to process audio data including spoken utterances to generate a transcription of the user's utterance (i.e., a sequence of terms and / or other tokens). The ASR engine 122A1 can utilize the on-device ASR model 132A1 to process the audio data. The on-device ASR model 132A1 can include, for example, a two-pass ASR model, which is a neural network model and is used by the ASR engine 122A1 to generate a sequence of probabilities over tokens (and the probabilities are used to generate the transcription). As another example, the on-device ASR model 132A1 can include an acoustic model that is a neural network model and a language model that includes a mapping from a sequence of phonemes to words. The ASR model 122A1 can use the acoustic model to process the audio data to generate a sequence of phonemes and use the language model to map the sequence of phonemes to specific terms. Additional or alternative ASR models can be utilized.
[0032] In Figure 1AAmong them, the assistant client 120A further includes a natural language understanding (NLU) engine 123A1 and one or more associated on-device NLU models 133A1. The NLU engine 123A1 can generate a symbolic representation or belief state that is a semantic representation of natural language text, such as the transcribed text generated by the ASR engine 122A1 or the typed text (e.g., typed using the virtual keyboard of the assistant device 110A). For example, the belief state can include the intent corresponding to the text and optionally include the parameters of the intent (e.g., slot values). Once fully formed through one or more conversation turns (e.g., all mandatory parameters have been parsed), the belief state representation represents the action to be performed in response to the spoken utterance. When generating the symbolic representation, the NLU engine 123A1 can utilize one or more on-device NLU models 133A1. The NLU model 133A1 can include one or more neural network models that are trained to process text and generate an output that indicates the intent expressed by the text and / or an indication of which part(s) of the text correspond to which parameter(s) of the intent. The NLU model can additionally or alternatively include one or more models that include a mapping of text and / or templates to corresponding symbolic representations. For example, the mapping can include a mapping of the text "what time is it" to the intent "current time" with the parameter "current location". As another example, the mapping can include a mapping of the template "add [item(s)] to my shopping list" to the intent "insert in shopping list", which intent has a parameter of the item(s) included in the actual natural language corresponding to [item(s)] in the template.
[0033] In Figure 1AIn this case, the assistant client 120A further includes an execution engine 124A1 and one or more associated on-device execution models 134A1. The execution engine 124A1 can utilize the fully formed symbolic representation from the NLU engine 123A1 to perform or cause the performance of actions corresponding to the symbolic representation. The actions can include providing response user interface outputs (such as audible and / or visual user interface outputs), controlling smart devices, and / or performing other actions. When performing or causing the performance of actions, the execution engine 124A1 can utilize the execution model 134A1. As an example, for the intent "turn on" with parameters specifying a particular smart light, the execution engine 124A1 can utilize the execution model 134A1 to identify the network address of the particular smart light and / or the command to be transmitted to cause the particular smart light to transition to the "on" state. As another example, for the intent "current" with the parameter "current location", the execution engine 124A1 can utilize the execution model 134A1 to identify that the current time at the client device 110A should be retrieved and audibly rendered (using the TTS engine 125A1).
[0034] In Figure 1A this case, the assistant client 120A further includes a text-to-speech (TTS) engine 125A1 and one or more associated on-device TTS models 135A1. The TTS engine 125A1 can utilize the on-device TTS model 135A1 to process text (or its speech representation) to generate synthetic speech. The synthetic speech can be audibly rendered via the speaker of the assistant device 110A's local text-to-speech ("TTS") engine (which converts text to speech). The synthetic speech can be generated and rendered as all or part of a response from the automated assistant and / or generated and rendered when prompting the user to define and / or clarify parameters and / or intents (such as orchestrated by the NLU engine 123A1 and / or a separate dialogue state engine).
[0035] In Figure 1AIn it, the assistant client 120A further includes an authentication engine 126A1 and one or more associated on-device authentication models 136A1. The authentication engine 126A1 can utilize one or more authentication techniques to verify which of the multiple registered users is interacting with the assistant device 110, or if only a single user is registered for the assistant device 110, whether it is the registered user interacting with the assistant device 110 (or alternatively a guest / unregistered user). As an example, with the permission of the associated user, text-dependent speaker verification (TD-SV) can be generated and stored for each of the registered users (e.g., associated with their corresponding user profiles). The authentication engine 126A1 can utilize the TD-SV model in the on-device authentication model 136A1 to generate the corresponding TD-SV and / or process the corresponding part of the audio data for TD-SV to generate a corresponding current TD-SV that can then be compared with the stored TD-SV to determine if there is a match. As other examples, the authentication engine 126A1 can additionally or alternatively utilize text-independent speaker verification (TI-SV) techniques, speaker verification techniques, facial verification techniques, and / or other verification techniques (e.g., PIN entry), and utilize the corresponding on-device authentication model 136A1 to authenticate a specific user.
[0036] In Figure 1A it, the assistant client 120A further includes a warm lead engine 127A1 and one or more associated on-device warm lead models 137A1. The warm lead engine 127A1 can at least selectively monitor the occurrence of one or more warm words or other warm leads, and in response to detecting one or more of the warm leads, cause a specific action to be performed by the assistant client 120A. The warm leads can be supplementary to any wake words or other wake-up leads, and each of the warm leads can be at least selectively active. Notably, detecting the occurrence of a warm lead causes the specific action to be performed even if the detected occurrence does not precede any wake-up leads. Thus, when the warm lead is one or more specific words, the user can simply say the words without the need to provide any wake-up leads, and cause the corresponding specific action to be performed.
[0037] As an example, a "stop" warm lead may be active at least when a timer or an alert is audibly rendered at the assistant device 110A via the automation assistant 120A. For example, at such a time, the warm lead engine 127A may continue (or at least when the VAD engine 128A1 detects voice activity) to process the audio data frame stream, which is based on the output from one or more microphones of the client device 110A, to monitor for the occurrence of "stop", "abort", or other limited set of specific warm words. The processing may be performed by the warm lead engine 127A using one of the warm lead models in the warm lead model 137A1, such as a neural network model trained to process audio data frames and generate an output indicating whether the spoken occurrence of "stop" exists in the audio data. In response to detecting the occurrence of "stop", the warm lead engine 127A may cause a command to clear the audible timer or alert to be implemented. At such a time, the warm lead engine 127A may continue (or at least when a sensor detects the presence) to process the image stream from the camera of the assistant device 110A to monitor for the occurrence of a hand in a "stop" pose. The processing may be performed by the warm lead engine 127A using one of the warm lead models in the warm lead model 137A1, such as a neural network model trained to process visual data frames and generate an output indicating whether a hand is present and in a "stop" pose. In response to detecting the occurrence of a "stop" pose, the warm lead engine 127A may cause a command to clear the audible timer or alert to be implemented.
[0038] As another example, a "volume up", "volume down", or "next" warm lead may be active at least when music is audibly rendered at the assistant device 110A via the automation assistant 120A. For example, at such a time, the warm lead engine 127A may continue to process the audio data frame stream, which is based on the output from one or more microphones of the client device 110A. The processing may include using a first warm lead model in the warm lead model 137A1 to monitor for the occurrence of "volume up", using a second warm lead model in the warm lead model 137A1 to monitor for the occurrence of "volume down", and using a third warm lead model in the warm lead model 137A1 to monitor for the occurrence of "next". In response to detecting the occurrence of "volume up", the warm lead engine 127A may cause a command to increase the volume of the music to be rendered to be implemented, in response to detecting the occurrence of "volume down", the warm lead engine may cause a command to decrease the music volume to be implemented, and in response to detecting the occurrence of "volume down", the warm lead engine may cause a command to cause the next track rather than the current music track to be rendered to be implemented.
[0039] At Figure 1AIn [the context], the assistant client 120A further includes a voice activity detector (VAD) engine 128A1 and one or more associated on-device VAD models 138A1. The VAD engine 128A1 can at least selectively monitor the occurrence of voice activity in audio data, and in response to detecting the occurrence, cause one or more functions to be performed by the assistant client 120A. For example, in response to detecting voice activity, the VAD engine 128A1 can cause the warm lead engine 121A1 to be activated. As another example, the VAD engine 128A1 can be used in a continuous listening mode to monitor the occurrence of voice activity in audio data, and in response to detecting the occurrence, cause the ASR engine 122A1 to be activated. The VAD engine 128A1 can utilize the VAD model 138A1 to process the audio data to determine whether voice activity is present in the audio data.
[0040] Specific engines and corresponding models have been described with respect to the assistant client 120A. However, it should be noted that some engines can be omitted and / or additional engines can be included. It should also be noted that through its various on-device engines and corresponding models, the assistant client 120A can fully process many assistant requests, including many assistant requests provided as spoken utterances. However, since the client device 110A is relatively limited in processing power, there are still many assistant requests that cannot be fully processed locally at the assistant device 110A. For example, the NLU engine 123A1 and / or the corresponding NLU model 133A1 may only cover a subset of all available intents and / or parameters available via the automated assistant. As another example, the fulfillment engine 124A1 and / or the corresponding fulfillment model may only cover a subset of available fulfillments. As yet another example, the ASR engine 122A1 and the corresponding ASR model 132A1 may not be robust and / or accurate enough to correctly transcribe various spoken utterances.
[0041] In view of these and other considerations, the cloud-based assistant component 140 can still be used at least selectively to perform at least some processing of assistant requests received at the assistant device 110A. The cloud-based automated assistant component 140 can include the corresponding (and / or additional or alternative) engines and / or models of these assistant devices 110A. However, since the cloud-based automated assistant component 140 can utilize the nearly infinite resources of the cloud, one or more of the cloud-based counterparts can be more robust and / or accurate with these assistant clients 120A. As an example, in response to an oral utterance seeking to perform an assistant action not supported by the local NLU engine 123A1 and / or the local fulfillment engine 124A1, the assistant client 120A can transmit the audio data of the oral utterance and / or its transcription generated by the ASR engine 122A1 to the cloud-based automated assistant component 140. The cloud-based automated assistant component 140 (e.g., its NLU engine and / or fulfillment engine) can perform more robust processing of such data, supporting the parsing and / or execution of the assistant action. Transmitting the data to the cloud-based automated assistant component 140 is via one or more wide area networks (WANs) 109, such as the Internet or a private WAN.
[0042] The second assistant device 110B includes an assistant client 120B, which can be a stand-alone application on top of the operating system or can form all or part of the operating system of the assistant device 110B. Similar to the assistant client 120A, the assistant client 120B includes: a wake / invoke engine 121B1 and one or more associated on-device wake / invoke models 131B1; an ASR engine 122B1 and one or more associated on-device ASR models 132B1; an NLU engine 123B1 and one or more associated on-device NLU models 133B1; a fulfillment engine 124B1 and one or more associated on-device fulfillment models 134B1; a TTS engine 125B1 and one or more associated on-device TTS models 135B1; an authentication engine 126B1 and one or more associated on-device authentication models 136B1; a warm lead engine 127B1 and one or more associated on-device warm lead models 137B1; and a VAD engine 128B1 and one or more associated on-device VAD models 138B1.
[0043] Some or all of the engines and / or models of assistant client 120B may be the same as and / or some or all of the engines and / or models may be different from those of assistant client 120A. For example, the wake cue engine 121B1 may lack the functionality to detect wake cues in images and / or the wake model 131B1 may lack the model to process images to detect wake cues, while the wake cue engine 121A1 includes such functionality and the wake model 131B1 includes such a model. For example, this may be because the assistant device 110A includes a camera and the assistant device 110B does not include a camera. As another example, the ASR model 131B1 utilized by the ASR engine 122B1 may be different from the ASR model 131A1 utilized by the ASR engine 122A1. For example, this may be because different models are optimized for different processor and / or memory capabilities in the assistant device 110A and the assistant device 110B.
[0044] Specific engines and corresponding models have been described with respect to assistant client 120B. However, it should be noted that some engines may be omitted and / or additional engines may be included. It should also be noted that, through its various on-device engines and corresponding models, the assistant client 120B can fully process many assistant requests, including many assistant requests provided as spoken utterances. However, since the client device 110B is relatively limited in processing capabilities, there are still many assistant requests that cannot be fully processed locally at the assistant device 110B. Given these and other considerations, the cloud-based assistant component 140 can still be at least selectively used to perform at least some processing of the assistant requests received at the assistant device 110B.
[0045] The third assistant device 110C includes an assistant client 120C, which may be a stand-alone application on top of the operating system or may form all or part of the operating system of the assistant device 110C. Similar to the assistant client 120A and the assistant client 120B, the assistant client 120C includes: a wake / invoke engine 121C1 and one or more associated on-device wake / invoke models 131C1; an authentication engine 126C1 and one or more associated on-device authentication models 136C1; a warm cue engine 127C1 and one or more associated on-device warm cue models 137C1; and a VAD engine 128C1 and one or more associated on-device VAD models 138C1. Some or all of the engines and / or models of the assistant client 120C may be the same as and / or some or all of the engines and / or models may be different from those of the assistant client 120A and / or the assistant client 120B.
[0046] However, it should be noted that, unlike assistant clients 120A and 120B, assistant client 120C does not include: any ASR engine or associated model; any NLU engine or associated model; any fulfillment engine or associated model; and any TTS engine or associated model. Further, it should also be noted that through its various on-device engines and corresponding models, assistant client 120B can only fully process specific assistant requests (i.e., assistant requests that match the warm leads detected by the warm lead engine 127C1), and cannot process many assistant requests, such as those provided as spoken utterances and that do not match the warm leads. Given these and other considerations, the cloud-based assistant component 140 can still be at least selectively used to perform at least some processing of the assistant requests received at the assistant device 110C.
[0047] The fourth assistant device 110D includes an assistant client 120D, which can be a stand-alone application on top of the operating system or can form all or part of the operating system of the assistant device 110D. Like assistant clients 120A, 120B, and 120C, assistant client 120D includes: a wake / invoke engine 121D1 and one or more associated on-device wake / invoke models 131D1; an authentication engine 126D1 and one or more associated on-device authentication models 136D1; a warm lead engine 127D1 and one or more associated on-device warm lead models 137D1; and a VAD engine 128D1 and one or more associated on-device VAD models 138D1. Some or all of the engines and / or models of assistant client 120C can be the same as and / or some or all of the engines and / or models of assistant clients 120A, 120B, and / or 120C can be different.
[0048] However, it should be noted that, unlike assistant clients 120A and 120B and like assistant client 120C, assistant client 120D does not include: any ASR engine or associated model; any NLU engine or associated model; any fulfillment engine or associated model; and any TTS engine or associated model. Further, it should also be noted that through its various on-device engines and corresponding models, assistant client 120D can only fully process specific assistant requests (i.e., assistant requests that match the warm leads detected by the warm lead engine 127D1), and cannot process many assistant requests, such as those provided as spoken utterances and that do not match the warm leads. Given these and other considerations, the cloud-based assistant component 140 can still be at least selectively used to perform at least some processing of the assistant requests received at the assistant device 110D.
[0049] Now turning to Figure 1B1 、 1B2, 1B3, 1C, and 1D illustrate different non - limiting examples of an assistant device group and different non - limiting examples of adaptations that can be implemented in response to the generation of the assistant device group. Through each of the adaptations, the grouped assistant devices can be collectively used to process various assistant requests, and through collective utilization, a more robust and / or accurate processing of these various assistant requests can be performed compared to any one of the assistant devices in the group that could be executed individually before the adaptation. This results in various technical advantages, such as the technical advantages described herein.
[0050] In Figure 1B1 , 1B2 , 1B3, 1C, and 1D, with respect to Figure 1A , the engines and models of the assistant clients with the same reference numerals as in Figure 1A are not adapted. For example, in Figure 1B1 , 1B2 and 1B3, the engines and models of the assistant client devices 110C and 110D are not adapted because the assistant client devices 110C and 110D are not included in the Figure 1B1 , 1B2 and 1B3's group 101B. However, in Figure 1B1 , 1B2 , 1B3, 1C, and 1D, the engines and models of the assistant clients with reference numerals different from Figure 1A (i.e., ending with "2", "3", or "4" instead of "1") indicate that it has been adapted relative to its counterpart in Figure 1A . Further, an engine or model with a reference numeral that ends with "2" in one figure and "3" in another figure means that a different adaptation of the engine or model has been made between the figures. Similarly, an engine or model with a reference numeral that ends with "4" in a figure means that the adaptation of the engine or model in that figure is different from the figures where the reference numerals end with "2" or "3".
[0051] Initially go to Figure 1B1, the device group 101B has been established, and the assistant devices 110A and 110B are included in the device group 101B. In some embodiments, the device group 101B can be generated in response to a user interface input that explicitly indicates the desire to group the assistant devices 110A and 110B. As an example, the user can provide the spoken utterance "group [label for assistant device 110A] and [label for assistant device 110B]" to any one of the assistant devices 110A to 110D. Such a spoken utterance can be processed by the corresponding assistant device and / or the cloud-based assistant component 140, interpreted as a request to group the assistant devices 110A and 110B, and the group 101B is generated in response to such an interpretation. As another example, a registered user of the assistant devices 110A to 110D can provide touch inputs at an application that supports configuring the settings of the assistant devices 110A to 110D. These touch inputs can explicitly specify that the assistant devices 110A and 110B will be grouped, and the group 101B can be generated in response. As yet another example, one of the example techniques described below for automatically generating the device group 101B can alternatively be used to determine that the device group 101B should be generated, but user input explicitly approving the generation of the device group 101B may be required before the device group 101B is generated. For example, a prompt indicating that the device group 101B should be generated can be rendered at one or more of the assistant devices 110A to 110D, and the device group 101B is actually generated only if an affirmative user interface input is received in response to the prompt (and optionally, if the user interface input is verified to be from a registered user).
[0052] In some embodiments, the device group 101B can alternatively be automatically generated. In some of these embodiments, the user interface output indicating the generation of the device group 101B can be rendered at one or more of the assistant devices 110A - 110D to inform the corresponding user of the group and / or registered users that they can override the automated generation of the device group 101B via user interface input. However, when the device group 101B is automatically generated, the device group 101B will be generated and corresponding adaptations will be made without first requesting a user interface input that explicitly indicates the desire to create a specific device group 101B (although earlier input can indicate general approval of creating the group). In some embodiments, the device group 101B can be automatically generated in response to determining that the assistant devices 110A and 110B meet one or more proximity conditions relative to each other. For example, the proximity conditions can include that the assistant devices 110A and 110B are assigned to the same structure (e.g., a specific home, a specific vacation home, a specific office) and / or the same room (e.g., kitchen, living room, dining room) or other area within the same structure in the device topology. As another example, the proximity conditions can include sensor signals from each of the assistant devices 110A and 110B indicating that they are in close proximity to each other. For example, if both of the assistant devices 110A and 110B continuously (e.g., greater than 70% of a time or other threshold) detect the occurrence of a wake word at the same time or close to the same time (e.g., within one second of each other), this can indicate that they are in close proximity to each other. Also, for example, one of the assistant devices 110A and 110B can emit a signal (e.g., ultrasound), and the other of the assistant devices 110A and 110B can attempt to detect the emitted signal. If the other of the assistant devices 110A and 110B detects the emitted signal, optionally with a threshold intensity, then it can indicate that they are in close proximity to each other. Additional and / or alternative techniques for determining temporal proximity and / or automatically generating device groups can be utilized.
[0053] Regardless of how the group 101B is generated, Figure 1B1Shows an example of the adaptation that can be performed on assistant devices 110A and 110B in response to them being included in group 101B. In various embodiments, one or both of assistant clients 120A and 120B can determine the adaptations that should be made and cause those adaptations to occur. In other embodiments, one or more engines of cloud-based assistant component 140 can additionally or alternatively determine the adaptations that should be made and cause those adaptations to occur. As described herein, the adaptations to be made can be determined based on taking into account the processing capabilities of both assistant clients 120A and 120B. For example, the adaptation can seek to utilize as much collective processing power as possible while ensuring that the individual processing capabilities of each assistant device in the assistant devices are sufficient for the engines and / or models to be stored and utilized locally at the assistant. Further, the adaptations to be made can also be determined based on usage data that reflects metrics related to the actual usage of the assistant devices in the group and / or other non-grouped assistant devices in the ecosystem. For example, if the processing power allows for either a larger but more accurate wake cue model or a larger but more robust warm cue model, but not both, the usage data can be used to choose between the two options. For example, if the usage data reflects the rare (or even non-existent) use of warm words and / or the detection of wake words generally barely exceeds the threshold and / or false negatives of wake words are commonly encountered, then a larger but more accurate wake cue model can be selected. On the other hand, if the usage data reflects the frequent use of warm words and / or the detection of wake words consistently exceeds the threshold and / or false negatives of wake words are rare, then a larger but more accurate warm cue model can be selected. Considering such usage data can be a factor in determining Figure 1B1 , 1B2 or the adaptation of 1B3 is selected because Figure 1B1 , 1B2 or 1B3 respectively show different adaptations of the same group 101B.
[0054] In Figure 1B1 , fulfillment engine 124A1 and fulfillment model 134A1, and TTS engine 125A1 and TTS model 135A1 have been removed from assistant device 110A. Further, assistant device 110A has a different ASR engine 122A2 and a different on-device ASR model 132A2, and a different NLU engine 123A2 and a different on-device NLU model 133A2. The different engines can be downloaded at assistant device 110A from local model repository 150, which is accessible via interaction with cloud-based assistant component 140. In Figure 1B1In [Assistant Device 110B], the wake-up cue engine 121B1, ASR engine 122B1, and authentication engine 126B1, and their corresponding models 131B1, 133B1, and 136B1 have been cleared from the assistant device 110B. Further, the assistant device 110B has a different NLU engine 123B2 and a different on-device NLU model 133B2, a different fulfillment engine 124B2 and fulfillment model 133B2, and a different warm word engine 127B2 and warm word model 137B2. The different engines can be downloaded at the assistant client 110B from the local model repository 150, which is accessible via interaction with the cloud-based assistant component 140.
[0055] Compared with the ASR engine 122A1 and ASR model 132A1, the ASR engine 122A2 and ASR model 132A2 of the assistant device 110A can be more robust and / or more accurate, but occupy more disk space, utilize more memory, and / or require more processor resources. For example, the ASR model 132A1 can include only a single-pass model, and the ASR model 132A2 can include a two-pass model.
[0056] Similarly, compared with the NLU engine 123A1 and NLU model 133A1, the NLU engine 123A2 and NLU model 133A2 can be more robust and / or more accurate, but occupy more disk space, utilize more memory, and / or require more processor resources. For example, the NLU model 133A1 can include only intents and parameters of a first classification, such as "lighting control", but the NLU model 133A2 can include intents of "illumination control" as well as "thermostat control", "smart lock control", and "reminder".
[0057] Therefore, the ASR engine 122A2, ASR model 132A2, NLU engine 123A2, and NLU model 133A2 are improved relative to their alternative counterparts. However, it should be noted that without first clearing the fulfillment engine 124A1 and fulfillment model 134A1, and the TTS engine 125A1 and TTS model 135A1, the processing power of the assistant device 110A may hinder the storage and / or use of the ASR engine 122A2, ASR model 132A2, NLU engine 123A2, and NLU model 133A2. Simply clearing such models from the assistant device 110A without complementary adaptation to the assistant device 110B and collaborative processing with the assistant device 110B will result in the assistant client 120A lacking the ability to fully locally process various assistant requests (i.e., without the need to utilize one or more cloud-based assistant components 140).
[0058] Accordingly, the assistant device 110B is supplemented and adapted, and collaborative processing between the assistant devices 110A and 110B occurs after the adaptation. Compared with the NLU engine 123B1 and the NLU model 133B1, the NLU engine 123B2 and the NLU model 133B2 of the assistant device 110B can be more robust and / or more accurate, but occupy more disk space, utilize more memory, and / or require more processor resources. For example, the NLU model 133B1 may include only the first-classified intents and parameters, such as "lighting control". However, the NLU model 133B2 can cover a larger number of intents and parameters. It should be noted that the intents and parameters covered by the NLU model 133B2 can be limited to the intents not yet covered by the NLU model 133A2 of the assistant client 120A. This can prevent duplication of functionality between the assistant clients 120A and 120B and expand the collective capabilities when the assistant clients 120A and 120B collaboratively process assistant requests.
[0059] Similarly, compared with the fulfillment engine 124B1 and the fulfillment model 124B1, the fulfillment engine 124B2 and the fulfillment model 134B2 can be more robust and / or more accurate, but occupy more disk space, utilize more memory, and / or require more processor resources. For example, the fulfillment model 124B1 may include only the fulfillment capabilities of a single classification of the NLU model 133B1, but the fulfillment model 124B2 can include the fulfillment capabilities of all classifications of the NLU model 133B2 as well as the NLU model 133A2.
[0060] Accordingly, the fulfillment engine 124B2, the fulfillment model 134B2, the NLU engine 123B2, and the NLU model 133B2 are improved compared to their alternative counterparts. However, without first clearing the cleared models and cleared engines from the assistant device 110B, the processing capabilities of the assistant device 110B may hinder the storage and / or use of the fulfillment engine 124B2, the fulfillment model 134B2, the NLU engine 123B2, and the NLU model 133B2. Simply clearing such models from the assistant device 110B without supplementary adaptation of the assistant device 110A and collaborative processing with the assistant device 110A will result in the assistant client 120B lacking the ability to fully locally process various assistant requests.
[0061] Compared with the warm lead engine 127B1 and the warm lead model 127B1, the warm lead engine 127B2 and the warm lead model 137B2 of the client device 110B do not occupy any additional disk space, utilize more memory, or require more processor resources. For example, they may require the same or even less processing power. However, the warm lead engine 127B2 and the warm lead model 137B2 cover warm leads that are complementary to the warm leads covered by the warm lead engine 127B1 and the warm lead model 127B1, and are also complementary to the warm leads covered by the warm lead engine 127A1 and the warm lead model 127A1 of the assistant client 120A.
[0062] In Figure 1B1 configuration, the assistant client 120A may be assigned the following processing roles: monitoring wake-up leads, performing ASR, performing NLU on a first set of classifications, performing authentication, monitoring a first set of warm leads, and performing VAD. The assistant client 120B may be assigned the following processing roles: performing NLU on a second set of classifications, performing fulfillment, performing TTS, and monitoring a second set of warm leads. The processing roles may be passed and stored at each of the assistant clients in the assistant client 120A, and the coordination of the processing of various assistant requests may be performed by one or both of the assistant clients 120A and 120B.
[0063] As used Figure 1B1An example of collaborative processing of an adapted assistant request is provided, assuming that the spoken utterance "OK Assistant, turn on the kitchen lights" is provided, and the assistant device 120A is the leading device for coordinated processing. The wake-up cue engine 121A1 of the assistant client 120A can detect the occurrence of the wake-up cue "OK Assistant". In response, the wake-up cue engine 121A1 can cause the ASR engine 122A2 to process the audio data captured after the wake-up cue. The wake-up cue engine 121A1 can also optionally transmit the command locally to the assistant device 110B to cause it to transition from a low-power state to a high-power state, so that the assistant client 120B is ready to perform some processing of the assistant request. The audio data processed by the ASR engine 122A2 can be the audio data captured by the microphone of the assistant device 110A and / or the audio data captured by the microphone of the assistant device 110B. For example, the command transmitted to the assistant device 110B to cause it to transition to a high-power state can also cause it to locally capture audio data and optionally transmit such audio data to the assistant client 120A. In some embodiments, based on an analysis of the characteristics of the corresponding instance of the audio data, the assistant client 120A can determine whether to use the received audio data or the locally captured audio data. For example, based on an instance with a lower signal-to-noise ratio and / or capturing a spoken utterance with a higher volume, an instance of the audio data can be utilized over another instance.
[0064] The transcription generated by the ASR engine 122A2 can be passed to the NLU engine 123A2 to perform NLU on the transcription of the first classification set, and also transmitted to the assistant client 120B to enable the NLU engine 123B2 to perform NLU on the transcription of the second classification set. The results of the NLU performed by the NLU engine 123B2 can be transmitted to the assistant client 120A, and it can determine which results (if any) to use based on these results and the results from the NLU engine 123A2. For example, the assistant client 120A can utilize the result with the highest probability intent as long as the probability meets some threshold. For example, a result including an intent of "turn on" and a parameter specifying an identifier of "kitchen light" can be utilized. It should be noted that if no probability meets the threshold, the NLU engine of the cloud-based assistant component 140 can be optionally used to perform NLU. The assistant client 120A can transmit the NLU result with the highest probability to the assistant client 120B. The fulfillment engine 124B2 of the assistant client 120B can utilize these NLU results to determine the command to be transmitted to the "kitchen light" to cause them to transition to the "on" state, and transmit such a command via the LAN 108. Optionally, the fulfillment engine 124B2 can utilize the TTS engine 125B1 to generate synthetic speech that confirms the execution of "turn on the kitchen light". In this case, the synthetic speech can be rendered by the assistant client 120B at the assistant device 110B and / or transmitted to the assistant device 110A to be rendered by the assistant client 120A.
[0065] As another example of collaborative processing for an assistant request, assume that assistant client 120A is rendering an alert for a just-expired local timer at assistant client 120A. Further assume that the warm lead monitored by warm lead engine 127B2 includes "stop", and the warm lead monitored by warm lead engine 127A1 does not include "stop". Finally, assume that while the alert is being rendered, the spoken utterance "stop" is provided and captured in the audio data detected via the microphone of assistant client 120B. The warm word engine 127B2 can process the audio data and determine that the word "stop" has occurred. Further, the warm word engine 127B2 can determine that the occurrence of the word stop is directly mapped to a command to clear the audible timer or alert. This command can be transmitted by assistant client 120B to assistant client 120A so that assistant client 120A can implement the command and clear the audible timer or alert. In some embodiments, the warm word engine 127B2 can monitor the occurrence of "stop" only in certain situations. In these embodiments, in response to rendering the alert or when an alert is expected to be rendered, assistant client 120A can transmit a command to cause the warm word engine 127B2 to monitor for the occurrence of "stop". The command can cause the monitoring to occur for a certain period of time or alternatively until an interruption of the monitoring command is sent.
[0066] Now turning to Figure 1B2 , the same group 101B is illustrated. In Figure 1B2 , the same adaptations as in Figure 1B1 have been made, except that ASR engine 121A1 and ASR model 132A1 are not replaced by ASR engine 122A2 and ASR model 132A2. Instead, ASR engine 121A1 and ASR model 132A1 are retained, and additional ASR engine 122A3 and additional ASR model 132A3 are provided.
[0067] ASR engine 121A1 and ASR model 132A1 can be used for speech recognition of utterances in a first language (e.g., English), and additional ASR engine 122A3 and additional ASR model 132A3 can be used for speech recognition of utterances in a second language (e.g., Spanish). Figure 1B1 The ASR engine 121A2 and ASR model 132A2 of
[0068] In Figure 1B2 the example, the local storage of the ASR engine 121A1 and the ASR model 132A1, as well as the additional ASR engine 122A3 and the additional ASR model 132A3, instead of Figure 1B1 the decision of the ASR engine 121A2 and the ASR model 132A2 can be based on using statistical information that indicates that the dictated utterances provided at the assistant devices 110A and 110B (and / or assistant devices 110C and 110D) include first language dictated utterances and second language dictated utterances. In Figure 1B1 the example, the statistical information can indicate only dictated utterances in the first language, resulting in the selection of Figure 1B1 the more robust ASR engine 121A2 and ASR model 132A2 in
[0069] Now turning to Figure 1B3 , the same group 101B is illustrated again. In Figure 1B3 it, the same adaptation as in Figure 1B1 has been made, except that: (1) the ASR engine 121A1 and the ASR model 132A1 are replaced by the ASR engine 122A4 and the ASR model 132A4, instead of being replaced by the ASR engine 122A2 and the ASR model 132A2; (2) there is no warm lead engine or warm lead model on the assistant device 110B; and (3) the ASR engine 122B4 and the ASR model 132B4 are locally stored and utilized on the assistant device 110B.
[0070] The ASR engine 122A4 and the ASR model 132A4 can be used to perform the first part of speech recognition, and the ASR engine 122B4 and the ASR model 132B4 can be used to perform the second part of speech recognition. For example, the ASR engine 122A4 can utilize the ASR model 132A4 to generate an output, which is transmitted to the assistant client 120B, and the ASR engine 122B4 can process the output when generating the recognition of the speech. As a specific example, the output can be a graph representing candidate recognitions, and the ASR engine 122B4 can perform a beam search on the graph when generating the recognition of the speech. As another specific example, the ASR model 132A4 can be the initial / downstream part (i.e., the first neural network layer) of an end-to-end speech recognition model, and the ASR model 132B4 can be the later / upstream part (i.e., the second neural network layer) of the end-to-end speech recognition model. In such an example, the end-to-end model is split between the two assistant devices 110A and 110B, and the output can be the state of the last layer (e.g., embedding) of the processed initial part. As yet another example, the ASR model 132A4 can be an acoustic model, and the ASR model 132B4 can be a language model. In such an example, the output can indicate a sequence of phonemes or a sequence of probability distributions of phonemes, and the ASR engine 122B4 can utilize the language model to select the transcription / recognition corresponding to the sequence.
[0071] The robustness and / or accuracy of the ASR engine 122A4, the ASR model 132A4, the ASR engine 122B4, and the ASR model 132B4 working collectively can exceed Figure 1B1 the robustness and / or accuracy of the ASR engine 122A2 and the ASR model 132A2. Further, the processing capabilities of the assistant clients 120A and 120B may prevent the ASR models 132A4 and 132B4 from being stored and utilized individually on either of the devices. However, the processing capabilities can support splitting the models and splitting the processing roles between the ASR engines 122A4 and 122B4 described herein. Note that on the assistant device 110B, clearing the warm lead engine and the warm lead model can support storing and utilizing the ASR engine 122B4 and the ASR model 132B4. In other words, the processing capabilities of the assistant device 110B cannot support storing and / or utilizing the warm lead engine and the warm lead model together with Figure 1B3 the other engines and models illustrated.
[0072] In Figure 1B3 the example of, locally storing the ASR engine 122A4, the ASR model 132A4, the ASR engine 122B4, and the ASR model 132B4 instead of Figure 1B1The decisions of the ASR engine 121A2 and the ASR model 132A2 can be based on the use of statistical information that indicates that speech recognition at the assistant devices 110A and 110B (and / or assistant devices 110C and 110D) is typically of low confidence and / or typically inaccurate. For example, the use statistical information can indicate that the confidence metric of the recognition is below the average (e.g., the average based on the user population) and / or the recognition is often corrected by the user (e.g., via editing the display of the transcription).
[0073] Now turning to Figure 1C , the device group 101C has been established, where the assistant devices 110A, 110B, and 110C are included in the device group 101C. In some embodiments, the device group 101C can be generated in response to a user interface input that explicitly indicates the desire to group the assistant devices 110A, 110B, and 110C. For example, the user interface input can indicate creating the device group 101C from scratch or alternatively adding the assistant device 110C to the device group 101B ( Figure 1B1 , 1B2 and 1B3) to create the desire to modify the group 101C. In some embodiments, the device group 101C can alternatively be automatically generated. For example, the device group 101B ( Figure 1B1 , 1B2 and 1B3) could previously have been generated based on determining that the assistant devices 110A and 110B are very close, and after creating the device group 101B, the assistant device 110C can be moved by the user such that it is now close to the devices 110A and 110B. Accordingly, the assistant device 110C can be automatically added to the device group 101B, thereby creating the modified group 101C.
[0074] Regardless of how the group 101C is generated, Figure 1C shows an example of the adaptation that can be made to the assistant devices 110A, 110B, and 110C in response to their inclusion in the group 101C.
[0075] In Figure 1C , the assistant device 110B already has the same adaptation as in Figure 1B3 . Further, the assistant device 110A has the same as in Figure 1B3The same adaptation in [Assistant Device 110A], except that: (1) the authentication engine 126A1 and the VAD engine 128A1 and their corresponding models 136A1 and 138A1 have been cleared; (2) the wake-up cue engine 121A1 and the wake-up cue model 131A1 have been replaced with the wake-up cue engine 121A2 and the wake-up cue model 131A2; and (3) there is no warm cue engine or warm cue model on the assistant device 110B. The models and engines stored on the assistant device 110C have not been adapted. However, the assistant client 120C can be adapted to implement cooperative processing of assistant requests with the assistant clients 120A and 120B.
[0076] In Figure 1CIn [the example], the authentication engine 126A1 and the VAD engine 128A1 have been removed from the assistant device 110A because counterparts already exist on the assistant device 110C. In some embodiments, the authentication engine 126A1 and / or the VAD engine 128A1 may only be removed after some or all of the data from these components has been merged with the counterparts already existing on the assistant device 110C. As an example, the authentication engine 126A1 may store voice embeddings of a first user and a second user, but the authentication engine 126C1 may only store the voice embeddings of the first user. Before removing the authentication engine 126A1, the voice embeddings of the second user may be locally transferred to the authentication engine 126C1 so that such voice embeddings can be utilized by the authentication engine 126C1, ensuring that the pre-adaptation capabilities are maintained after adaptation. As another example, the authentication engine 126A1 may store instances of audio data that respectively capture the utterances of the second user and are used to generate the voice embeddings of the second user, and the authentication engine 126C1 may lack any voice embeddings of the second user. Before removing the authentication engine 126A1, the instances of audio data may be locally transferred from the authentication engine 126A1 to the authentication engine 126C1 so that the instances of audio data can be utilized by the authentication engine 126C1 to generate voice embeddings for the second user, using the on-device authentication model 136C1, ensuring that the pre-adaptation capabilities are maintained after adaptation. Further, the wake cue engine 121A1 and the wake cue model 131A1 have been replaced by the wake cue engine 121A2 and the wake cue model 131A2 with a smaller storage size. For example, the wake cue engine 121A1 and the wake cue model 131A1 support detecting verbal wake cues and image-based wake cues, while the wake cue engine 121A2 and the wake cue model 131A2 only support detecting image-based wake cues. Optionally, before removing the wake cue engine 121A1 and the wake cue model 131A1, the personalization, training instances, and / or other settings of the image-based wake cue portion from the wake cue engine 121A1 and the wake cue model 131A1 may be merged with or otherwise shared with the wake cue engine 121A2 and the wake cue model 131A2. The wake cue engine 121C1 and the wake cue model 131C1 only support detecting verbal wake cues. Thus, the wake cue engine 121A2 and the wake cue model 131A2, together with the wake cue engine 121C1 and the wake cue model 131C1, collectively support detecting verbal and image-based wake cues. Optionally, the personalization and / or other settings of the verbal cue portion from the wake cue engine 121A1 and the wake cue model 131A1 may be transferred to the client device 110C to be merged with or otherwise shared with the wake cue engine 121C1 and the wake cue model 131C1.
[0077] Further, replacing the wake cue engine 121A1 and the wake cue model 131A1 with the wake cue engine 121A2 and the wake cue model 131A2 of smaller storage size provides additional storage space. This additional storage space, along with the additional storage space provided by clearing the authentication engine 126A1 and the VAD engine 128A1 and their corresponding models 13A1 and 138A1, provides space for the warm cue engine 127A2 and the warm cue model 137A2 (collectively larger than the warm cue engine 127A1 and the warm cue model they replace). The warm cue engine 127A2 and the warm cue model 137A2 can be used to monitor different warm cues than those monitored by the warm cue engine 127C1 and the warm cue model 137C1.
[0078] As an example of collaborative processing of an adapted assistant request using Figure 1C assume that the spoken utterance "OK Assistant, turn on the kitchen lights" is provided and the assistant device 120A is the lead device for coordinated processing. The wake cue engine 121C1 of the assistant client 110C can detect the occurrence of the wake cue "OK Assistant". In response, the wake cue engine 121C1 can transmit a command to the assistant devices 110A and 110B to have the ASR engines 122A4 and 122B4 collaboratively process the audio data captured after the wake cue. The processed audio data can be the audio data captured by the microphone of the assistant device 110C and / or the audio data captured by the microphones of the assistant devices 110B and / or 110C.
[0079] The transcription generated by the ASR engine 122B4 can be passed to the NLU engine 123B2 to perform NLU on the transcription of the second classification set, and also transmitted to the assistant client 120A to enable the NLU engine 123A2 to perform NLU on the transcription of the first classification set. The results of the NLU performed by the NLU engine 123B2 can be transmitted to the assistant client 120A, and it can determine which results (if any) to use based on these results and the results from the NLU engine 123A2. The assistant client 120A can transmit the NLU result with the highest probability to the assistant client 120B. The fulfillment engine 124B2 of the assistant client 120B can utilize these NLU results to determine the command to be transmitted to the "kitchen lights" to cause them to transition to the "on" state, and transmit such a command via the LAN 108. Optionally, the fulfillment engine 124B2 can utilize the TTS engine 125B1 to generate synthetic speech that confirms the execution of "turn on the kitchen lights". In this case, the synthetic speech can be rendered by the assistant client 120B at the assistant device 110B, transmitted to the assistant device 110A for rendering by the assistant client 120A, and / or transmitted to the assistant device 110C for rendering by the assistant client 120C.
[0080] Now turning to Figure 1D , the device group 101D has been established, where the assistant devices 110C and 110D are included in the device group 101D. In some embodiments, the device group 101D can be generated in response to a user interface input that explicitly indicates the desire to group the assistant devices 110C and 110D. In some embodiments, the device group 101D can alternatively be generated automatically.
[0081] Regardless of how the group 101D is generated, Figure 1D an example of the adaptation that can be performed on the assistant devices 110C and 110D in response to their inclusion in the group 101D is shown.
[0082] In Figure 1D , the wake cue engine 121C1 and wake cue model 131C1 of the assistant device 110C are replaced with the wake cue engine 121C2 and wake cue model 131C2. Further, the authentication engine 126D1 and authentication model 136D1, as well as the VAD engine 128D1 and VAD model 138D2, are removed from the assistant device 110D. Furthermore, the wake cue engine 121D1 and wake cue model 131D1 of the assistant device 110D are replaced with the wake cue engine 121D2 and wake cue model 131D2, and the warm cue engine 127D1 and warm cue model 137D1 are replaced with the warm cue engine 127D2 and warm cue model 137D2.
[0083] The previous wake cue engine 121C1 and wake cue model 131C1 can be used only to detect a first set of one or more wake words, such as "Hey Assistant" and "OK Assistant". On the other hand, the wake cue engine 121C2 and wake cue model 131C2 can detect only an alternate second set of one or more wake words, such as "Hey Computer" and "OK Computer". The previous wake cue engine 121D1 and wake cue model 131D1 can also be used only to detect the first set of one or more wake words, and the wake cue engine 121D2 and wake cue model 131D2 can also be used only to detect the first set of one or more wake words. However, the wake cue engine 121D2 and wake cue model 131D2 are larger than their replaced counterparts and are also more robust (e.g., more robust to background noise) and / or more accurate. Clearing the engines and models from the assistant device 110D can enable the use of the larger-sized wake cue engine 121D2 and wake cue model 131D2. Further, collectively, the wake cue engine 121C2 and wake cue model 131C2 and the wake cue engine 121D2 and wake cue model 131D2 support the detection of two sets of wake words, while each of the assistant clients 120C and 120D can detect only the first set before adaptation.
[0084] Compared with the replaced wake cue engine 127D1 and wake cue model 137D1, the warm cue engine 127D2 and warm cue model 137D2 of the assistant device 110D may require more computing power. However, these capabilities are available by clearing the engines and models from the assistant device 110D. Also, the warm cues monitored by the warm cue engine 127D2 and warm cue model 137D1 can be complementary to the warm cues monitored by the warm cue engine 127C1 and warm cue model 137D1. Before adaptation, the wake cues monitored by the wake cue engine 127D1 and wake cue model 137D1 are the same as the warm cues monitored by the warm cue engine 127C1 and warm cue model 137D1. Thus, through collaborative processing, the assistant client 120C and the assistant client 120D can monitor a greater number of wake cues.
[0085] Note that in Figure 1DIn the example of, there are many assistant requests that cannot be fully processed collaboratively by assistant clients 120C and 120D on the device. For example, assistant clients 120C and 120D lack any ASR engine, any NLU engine, and any fulfillment engine. This may be due to the processing capabilities of assistant devices 110C and 110D not supporting any such engine or model. Therefore, for the dictated utterances of warm leads not supported by assistant clients 120C and 120D, the cloud-based assistant component 140 will still be needed to fully process many assistant requests. However, compared to any processing that occurred individually at the device before adaptation, Figure 1D the adaptation and the collaborative processing based on the adaptation can still be more robust and / or accurate. For example, the adaptation supports detecting additional wake-up cues and additional warm cues.
[0086] As an example of the possible collaborative processing, assume the dictated utterance "OK Computer, play some music". In this example, the wake-up cue engine 121D2 can detect the wake-up cue "OK Computer". In response, the wake-up cue engine 121D2 can cause the audio data corresponding to the wake-up cue to be transmitted to assistant client 120C. The authentication engine 126C1 of assistant client 120C can use the audio data to determine whether the dictated wake-up cue is authenticated to a registered user. The wake-up cue engine 121D2 can also cause the audio data after the dictated utterance to be streamed to the cloud-based assistant component 140 for further processing. The audio data can be captured at assistant device 110D or at assistant device 110C (e.g., assistant client 120 can transmit a command to assistant client 120C to capture the audio data in response to the wake-up cue engine 121D2 detecting the wake-up cue). Further, the authentication data based on the output of the authentication engine 126C1 can also be transmitted together with the audio data. For example, if the authentication engine 126C1 authenticates the dictated wake-up cue to a registered user, the authentication data can include the identifier of the registered user. As another example, if the authentication engine 126C1 does not authenticate the dictated wake-up cue to any registered user, the authentication data can include an identifier reflecting that the utterance was provided by a guest user.
[0087] Various specific examples have been described with reference to Figure 1B1 、 1B2 、1B3, 1C, and 1D. However, it should be noted that various additional or alternative groups can be generated and / or various additional or alternative adaptations can be performed in response to the generation of the groups.
[0088] Figure 2FIG. 200 is a flow chart of an example method for diagramming the device - side models and / or processing roles of assistant devices in an adaptation group. For convenience, the operations of the flow chart are described with reference to the system that performs the operations. The system may include various components of various computer systems, such as Figure 1A , Figure 1B1 , Figure 1B2 , Figure 1B3 , Figure 1C and Figure 1D automated clients 120A - 120D of Figure 1A , Figure 1B1 , Figure 1B2 , Figure 1B3 , Figure 1C and Figure 1D and / or one or more of the components of the cloud - based assistant components of Figure 1A , Figure 1B1 , Figure 1B2 , Figure 1B3 , Figure 1C and Figure 1D . Moreover, although the operations of method 200 are shown in a particular order, this is not intended as a limitation. One or more operations may be reordered, omitted, or added.
[0089] In block 252, the system generates a group of assistant devices. For example, the system may generate a group of assistant devices in response to a user interface input that explicitly indicates the desire to generate the group. As another example, the system may automatically generate the group in response to determining that one or more conditions are met. As yet another example, the system may automatically determine that the group should be generated in response to determining that a condition is met, provide a user interface output that recommends generating the group, and then generate the group in response to an affirmative user interface received in response to the user interface output.
[0090] In block 254, the system obtains the processing capabilities of each assistant device in the group of assistant devices. For example, the system may be one of the assistant devices in the group of assistant devices. In such an example, the assistant device may obtain its own processing capability, and the other assistant devices in the group may transmit their processing capabilities to the assistant device. As another example, the processing capabilities of the assistant devices may be stored in a device topology, and the system may retrieve them from the device topology. As yet another example, the system may be a cloud - based component, and the assistant devices in the group may transmit their processing capabilities to the system, respectively.
[0091] The processing capabilities of the assistant device can include corresponding processor values based on the capabilities of processors on one or more devices, corresponding memory values based on the size of the memory on the device, and / or corresponding disk space values based on the available disk space. For example, the processor values can include details about one or more operating frequencies of the processor, details about the size of the cache of the processor, whether each processor in the processor is a GPU, CPU, or DSP, and / or other details. As another example, the processor values can additionally or alternatively include a higher-level categorization of the capabilities of the processor, such as high, medium, or low, or GPU+CPU+DSP, high-power CPU+DSP, medium-power CPU+DSP, or low-power CPU+DSP. As another example, the memory values can include details of the memory (such as the specific size of the memory), or can include a higher-level categorization of the memory, such as high, medium, or low. As yet another example, the disk space values can include details about the available disk space (such as the specific size of the disk space), or can include a higher-level categorization of the available disk space, such as high, medium, or low.
[0092] In block 256, the system utilizes the processing capabilities of block 254 when determining the collective set of on-device models for the group. For example, the system can determine a set of on-device models that seeks to maximize the use of collective processing capabilities while ensuring that the on-device models in the set can be locally stored and used respectively on the device that can store and use the on-device models. The system can also seek to ensure (if possible) that the selected set includes a complete (or more complete than other candidate sets) pipeline of on-device models. For example, compared to a set that includes a more robust NLU model but no ASR model, the system can select a set that includes an ASR model but the NLU model is less robust.
[0093] In some embodiments, block 256 includes sub - block 256A, where the system utilizes usage data when selecting the overall set of device - on models for the group. Past usage data can be data related to past assistant interactions at one or more of the assistant devices in the group and / or at one or more additional assistant devices of the ecosystem. In some embodiments, in sub - block 256A, the system considers the usage data and the considerations mentioned above when selecting device - on models to include in the set. For example, if processing power allows a more accurate ASR model (as opposed to a less accurate ASR model) or a more robust NLU model (as opposed to a less robust NLU model) to be included in the set but not both, the usage data can be used to determine which one to select. For example, if the usage data reflects that past assistant interactions mainly (or exclusively) involve intents covered by a less robust NLU model (which can be included in a set with a more accurate ASR model), then the more accurate ASR model can be selected to be included in the set. On the other hand, if the usage data reflects that many of the past assistant interactions included involve intents covered by a more robust NLU model rather than a less robust NLU model, then the more robust NLU model can be selected to be included in the set. In some embodiments, a candidate set is first determined based on processing power without considering usage data, and then, if there are multiple valid candidate sets, the usage data can be used to select one model over the others.
[0094] In block 258, the system causes each assistant device in the assistant device group to locally store a corresponding subset of the overall set of models on the device. For example, the system can pass to each assistant device in the group a corresponding indication of which on-device models should be downloaded. Based on the received indication, each assistant device in the assistant device group can then download the corresponding model from the remote database. As another example, the system can retrieve the on-device models and push the corresponding on-device models to each assistant device in the group. As yet another example, for any on-device model that was stored at a corresponding assistant device in the assistant device group before adaptation and will be stored at another corresponding assistant device in the assistant device group during adaptation, such a model can be directly transferred between the respective devices. For example, assume that a first assistant device stores an ASR model before adaptation, and during adaptation, the same ASR model will be stored on a second assistant device and cleared from the first assistant device. In such an instance, the system can instruct the first assistant device to transfer the ASR model to the second assistant device for local storage at the second assistant device (and / or instruct the second assistant device to download it from the first assistant device), and the first assistant device can then clear the ASR model. In addition to avoiding WAN traffic, locally transferring the pre-adaptation models can maintain any personalization that previously occurred for these on-device models during transfer. For users of the ecosystem, a personalized model can be more accurate compared to the non-personalized counterpart of the model in the remote storage device. As yet another example, for those assistant devices that included stored training instances before adaptation to personalize the on-device models on any assistant device before adaptation, such training instances can be passed to the assistant devices that will have the corresponding models downloaded from the remote database after adaptation. The assistant devices that have the on-device models after adaptation can then use the training instances to personalize the corresponding models downloaded from the remote database. The corresponding models downloaded from the remote database can be different from (e.g., smaller or larger than) the counterparts on which the training instances were utilized before adaptation, but the training instances can still be used to personalize the different downloaded on-device models.
[0095] In block 260, the system assigns corresponding roles to each of the assistant devices in the assistant device group. In some embodiments, assigning the corresponding roles includes causing each of the assistant devices in the assistant device group to download and / or implement an engine corresponding to a device-on model locally stored at the assistant device. The engine can utilize the corresponding device-on model when performing corresponding processing roles, such as performing all or part of the ASR, performing wake word recognition for at least some wake words, performing warm word recognition for certain warm words, and / or performing authentication. In some embodiments, one or more of the processing roles are performed only upon the command of a master device in the assistant device group. For example, the NLU processing role performed by a given device using a device-on NLU model can be performed only in response to the master device transmitting corresponding text for NLU processing and / or a specific command causing the NLU processing to occur to the given device. As another example, the warm word monitoring processing role performed by a given device using a device-on warm lead engine and a device-on warm lead model can be performed only in response to the master device transmitting a command causing the warm word processing to occur to the given device. For example, the master device can cause a given device to monitor for the spoken occurrence of the "stop" warm word in response to an alarm sounding at the master device or another device in the group. In some embodiments, one or more of the processing roles can be performed at least selectively independently of any command from the master assistant device. For example, the warm lead monitoring role performed by a given device using a device-on warm lead engine and a device-on warm lead model can be performed continuously unless explicitly disabled by the user. As another example, the warm lead monitoring role performed by a given device using a device-on warm lead engine and a device-on warm lead model can be performed continuously or based on monitoring conditions locally detected at the given device.
[0096] In block 262, the system causes spoken utterances detected at one or more of the devices in the group to be collaboratively and locally processed at the assistant devices in the group according to their roles. Various non-limiting examples of such collaborative processing are described herein. For example, examples are referenced Figure 1B1 、 1B2 、1B3, 1C, and 1D.
[0097] In block 264, the system determines whether there are any changes to the group, such as adding a device to the group, removing a device from the group, or disabling the group. If not, the system continues to perform block 262. If so, the system proceeds to block 266.
[0098] In block 266, the system determines whether the change to the group has caused one or more of the assistant devices in the group to now be solo (i.e., no longer assigned to the group). If so, the system proceeds to block 268 and causes each of the solo devices in the solo devices to locally store the model on the pre-group device and assume the processing role on the pre-group device. In other words, if the device is no longer in the group, it can be caused to revert to the state before the adaptation performed in response to its inclusion in the group. In these and other ways, after reverting to this state, the solo device can operate in a solo capacity to functionally process various assistant requests. Before reverting to this state, the solo device may not be able to functionally process any assistant requests or at least fewer assistant requests compared to the assistant requests it could process before reverting to this state.
[0099] In block 270, the system determines whether two or more devices remain in the changed group. If so, the system returns to block 254 and performs another iteration of blocks 254, 256, 258, 260, and 262 based on the changed group. For example, if the changed group includes additional assistant devices without losing any of the previous assistant devices in the group, the adaptation can be made taking into account the additional processing capabilities of the additional assistant devices. If the decision in block 270 is no, the group has been disbanded and the system proceeds to block 272 where method 200 ends (until another group is generated).
[0100] Figure 3 FIG. is a flow chart of an example method 300 that can be implemented by each of multiple assistant devices in a group when adapting a device model and / or processing role of an assistant device in an adaptation group. Although the operations of method 300 are shown in a particular order, this is not intended to be limiting. One or more operations can be reordered, omitted, or added.
[0101] The operations of method 300 are a particular example of method 200 that can be performed by each of the assistant devices in the group. Thus, the operations are described with reference to the assistant devices (such as one or more of assistant clients 120A through 120D of Figure 1A , Figure 1B1 , Figure 1B2 , Figure 1B3 , Figure 1C and Figure 1D ) that perform the operations. Each of the assistant devices in the group can perform method 300 in response to receiving an input indicating that it has been included in the group.
[0102] In block 352, the assistant device receives a grouping indication that it has been included in a group. In block 352, the assistant device also receives identifiers of other assistant devices in the group. For example, each identifier can be a MAC address, an IP address, a label assigned to the device (such as for use assigned in a device topology), a serial number, or other identifier.
[0103] In optional block 354, the assistant device transmits data to other assistant devices in the group. The data is transmitted to the other devices using the identifiers received in block 352. In other words, the identifier can be a network address or can be used to find the network address to which the data is to be transmitted. The transmitted data can include one or more processing values described herein, another device identifier, and / or other data.
[0104] In optional block 356, the assistant device receives data transmitted by other devices in block 354.
[0105] In block 358, based on the data optionally received in block 356 or the identifiers received in block 352, the assistant device determines whether it is the leading device. For example, if the device's own identifier is the lowest (alternatively, the highest) value compared to the other identifiers received in block 352, the device can select itself as the leader. As another example, if the device's processing value exceeds all other processing values received in the data in optional block 356, the device can select itself as the leader. Other data can be transmitted in block 354 and received in block 356, and such other data can similarly enable an objective determination at the assistant device as to whether it should be the leader. More generally, in block 358, the assistant device can utilize one or more objective criteria when determining whether it should be the leading device.
[0106] In block 360, the assistant device determines whether it has been determined to be the leading device in block 358. Assistant devices not determined to be the leading device will then proceed to the "No" branch of block 360. Assistant devices determined to be the leading device will proceed to the "Yes" branch of block 360.
[0107] In the "Yes" branch, in block 360, the assistant device utilizes the processing capabilities received from other assistant devices in the group and its own processing capabilities when determining the overall set of device models in the group. When optional block 354 is not executed or the data in block 354 does not include processing capabilities, the processing capabilities can be transmitted to the leading device by other assistant devices in optional block 352 or in block 370 (described below). In some embodiments, block 362 can be associated with Figure 2The blocks 256 of method 200 share one or more common aspects. For example, in some embodiments, block 362 may also include considering past usage data when determining the overall set of device models.
[0108] In block 364, the assistant device transmits to each of the other assistant devices in the group a corresponding indication of the device model on the overall set that the other assistant devices are to download. In block 364, the assistant device may also optionally transmit to each of the other assistant devices in the group a corresponding indication of the processing role that will be performed by that assistant device using the device model.
[0109] In block 366, the assistant device downloads and stores the device models in the set that are assigned to that assistant device. In some embodiments, blocks 364 and 366 may share one or more common aspects with Figure 2 block 258 of method 200.
[0110] In block 368, the assistant device coordinates the collaborative processing of the assistant request, including using its own device model when performing a part of the collaborative processing. In some embodiments, block 368 may share one or more common aspects with Figure 2 block 262 of method 200.
[0111] Now turning to the "No" branch, in optional block 370, the assistant device transmits its processing capabilities to the dominant device. For example, when block 354 is executed and the processing capabilities are included in the data transmitted in block 354, block 370 may be omitted.
[0112] In block 372, the assistant device receives from the dominant device an indication of the device model to be downloaded and optionally an indication of the processing role.
[0113] In block 374, the assistant device downloads and stores the device models reflected in the indication of the device model received in block 372. In some embodiments, blocks 372 and 374 may share one or more common aspects with Figure 2 block 258 of method 200.
[0114] In block 376, the assistant device uses its device model when performing a part of the collaborative processing of the assistant request. In some embodiments, block 376 may share one or more common aspects with Figure 2 block 262 of method 200.
[0115] Figure 4FIG. 0 is a block diagram of an example computing device 410 that can be optionally used to perform one or more aspects of the techniques described herein. In some embodiments, one or more of the assistant device and / or other components may include one or more components of the example computing device 410.
[0116] Computing device 410 generally includes at least one processor 414 that communicates with a number of peripheral devices via bus subsystem 412. These peripheral devices may include storage subsystem 425 (including, for example, memory subsystem 425 and file storage subsystem 426), user interface output device 420, user interface input device 422, and network interface subsystem 416. The input and output devices allow a user to interact with computing device 410. Network interface subsystem 416 provides an interface to external networks and is coupled to corresponding interface devices in other computing devices.
[0117] User interface input device 422 may include a keyboard, pointing device (such as a mouse, trackball, touchpad, or graphics tablet), scanner, touchscreen incorporated into a display, audio input device (such as a voice recognition system, microphone), and / or other types of input devices. In general, the use of the term "input device" is intended to include all possible types of devices and ways of inputting information into computing device 410 or into a communication network.
[0118] User interface output device 420 may include a display subsystem, printer, fax machine, or non-visual display (such as an audio output device). The display subsystem may include a cathode ray tube ("CRT"), flat panel device (such as a liquid crystal display ("LCD")), projection device, or some other mechanism for creating a visible image. The display subsystem may also provide a non-visual display, such as via an audio output device. In general, the use of the term "output device" is intended to include all possible types of devices and ways of outputting information from computing device 410 to a user or to another machine or computing device.
[0119] Storage subsystem 425 stores programming and data constructs that provide some or all of the functionality of the modules described herein. For example, storage subsystem 425 may include logic for performing selected aspects of one or more of the methods described herein and / or implementing the various components depicted herein.
[0120] These software modules are typically executed by the processor 414 alone or in combination with other processors. The memory 425 used in the storage subsystem 425 may include several memories, including a main random access memory ("RAM") 430 for storing instructions and data during program execution and a read-only memory ("ROM") 432 in which fixed instructions are stored. The file storage subsystem 426 may provide persistent storage for program and data files and may include a hard disk drive, a floppy disk drive with an associated removable medium, a CD-ROM drive, an optical drive, or a removable media cartridge. Modules implementing the functionality of a particular embodiment may be stored in the storage subsystem 425 by the file storage subsystem 426 or in other machines accessible by the processor 414.
[0121] The bus subsystem 412 provides the means for enabling the various components and subsystems of the computing device 410 to communicate with each other as intended. Although the bus subsystem 412 is schematically shown as a single bus, alternative embodiments of the bus subsystem may use multiple buses.
[0122] The computing device 410 may be of different types, including workstations, servers, computing clusters, blade servers, server farms, or any other data processing system or computing device. Due to the ever-changing nature of computers and networks, for the purpose of illustrating some embodiments, Figure 4 the description of the computing device 410 depicted herein is only intended as a specific example. Many other configurations of the computing device 410 with more or fewer components are possible compared to Figure 4 the computing device depicted herein.
[0123] In cases where the systems described herein collect personal information about a user (or generally referred to herein as a "participant") or may utilize personal information, the user may be provided with the following opportunities: to control whether a program or feature collects user information (such as information about the user's social network, social actions or activities, occupation, user preferences, or the user's current geographical location) or to control whether and / or how to receive content that may be more relevant to the user from a content server. Also, before a particular piece of data is stored or used, the particular data may be disposed of in one or more ways such that personally identifiable information is removed. For example, the user's identity may be disposed of such that the user's personally identifiable information cannot be determined, or the user's geographical location (from which geographical location information such as city, zip code, or state is obtained) may be generalized such that the user's specific geographical location cannot be determined. Thus, the user can control how information about the user is collected and / or used.
[0124] In some embodiments, a method is provided that includes: generating a group of assistant devices of different assistant devices. The different assistant devices include at least a first assistant device and a second assistant device. When generating the group, the first assistant device includes a first set of device - on - local - storage models used when locally processing assistant requests directed to the first assistant device. Further, when generating the group, the second assistant device includes a second set of device - on - local - storage models used when locally processing assistant requests directed to the second assistant device. The method further includes: determining a total set of device - on - local - storage models for collaboratively locally processing assistant requests directed to any one of the different assistant devices in the group of assistant devices, based on the corresponding processing capabilities of each of the different assistant devices in the group of assistant devices. The method also includes: in response to generating the group of assistant devices, causing each of the different assistant devices in the group of assistant devices to locally store a corresponding subset of the total set of device - on - local - storage models, and assigning one or more corresponding processing roles to each of the different assistant devices in the group of assistant devices. Each of the processing roles utilizes one or more of the corresponding device - on - local - storage models in the locally stored device - on - local - storage models. Further, causing each of the different assistant devices in the group of assistant devices to locally store the corresponding subset includes: causing the first assistant device to clear one or more of the first device - on - local - storage models in the first set to provide storage space for the corresponding subset locally stored on the first assistant device, and causing the second assistant device to clear one or more of the second device - on - local - storage models in the second set to provide storage space for the corresponding subset locally stored on the second assistant device. The method also includes: after assigning the corresponding processing roles to each of the different assistant devices in the group of assistant devices: detecting spoken utterances via a microphone of at least one of the different assistant devices in the group of assistant devices, and in response to the spoken utterances being detected via the microphone of the group of assistant devices, causing the spoken utterances to be collaboratively locally processed by the different assistant devices in the group of assistant devices using their corresponding processing roles.
[0125] These and other embodiments of the technology disclosed herein may optionally include one or more of the following features.
[0126] In some embodiments, causing a first assistant device to clear a model on one or more first devices in a first set includes: causing the first assistant device to clear a first device wake word detection model in the first set that is used to detect a first wake word. In these embodiments, a corresponding subset locally stored on a second assistant device includes a second device wake word detection model that is used to detect the first wake word, and assigning a corresponding processing role includes: assigning a first wake word detection role to the second assistant device, the first wake word detection role utilizing the second device wake word detection model when monitoring for the occurrence of the first wake word. In some of these embodiments, an uttered discourse includes the first wake word followed by an assistant command, and in the first wake word detection role, the second assistant device detects the occurrence of the first wake word and causes the execution of an additional processing role in the corresponding processing role in response to detecting the occurrence of the first wake word. In some versions of these embodiments, the additional processing role in the corresponding processing role is executed by the first assistant device, and the second assistant device causes the execution of the additional processing role in the corresponding processing role by transmitting an indication of the detection of the first wake word to the first assistant device.
[0127] In some embodiments, a corresponding subset locally stored on the first assistant device includes a first device first wake word detection model that is used to detect one or more first wake words and does not include any wake word detection models that are used to detect one or more second wake words. In some of these embodiments, a corresponding subset locally stored on the second assistant device includes a second device second wake word detection model that is used to detect one or more second hot words and does not include any wake word detection models that are used to detect one or more first wake words. In some versions of these embodiments, assigning a corresponding processing role includes: assigning a first wake word detection role to the first assistant device, the first wake word detection role utilizing the first device wake word detection model when monitoring for the occurrence of one or more first wake words; and assigning a second wake word detection role to the second assistant device, the second wake word detection role utilizing the second device wake word detection model when monitoring for the occurrence of one or more second wake words.
[0128] In some embodiments, the corresponding subset locally stored on the first assistant device includes a first language speech recognition model used to perform speech recognition of a first language and does not include any speech recognition models used to recognize speech in a second language. In some of these embodiments, the corresponding subset locally stored on the second assistant device includes a second language speech recognition model used to perform speech recognition of a second language and does not include any speech recognition models used to recognize speech in the second language. In some versions of these embodiments, assigning corresponding processing roles includes: assigning a first language speech recognition role to the first assistant device, the first language speech recognition role utilizing the first language speech recognition model when performing speech recognition of the first language; and assigning a second language speech recognition role to the second assistant device, the second language speech recognition role utilizing the second language speech recognition model when performing speech recognition of the second language.
[0129] In some embodiments, the corresponding subset locally stored on the first assistant device includes a first part of a speech recognition model used to perform a first part of speech recognition and does not include a second part of the speech recognition model. In some of these embodiments, the corresponding subset locally stored on the second assistant device includes a second part of a speech recognition model used to perform a second part of speech recognition and does not include the first part of the speech recognition model. In some versions of these embodiments, assigning corresponding processing roles includes: assigning a first part of a language speech recognition role to the first assistant device, the first part of the language speech recognition role utilizing the first part of the speech recognition model when generating a corresponding embedding of the corresponding speech and transmitting the corresponding embedding to the second assistant device; and assigning a second language speech recognition role to the second assistant device, the second language speech recognition role utilizing the corresponding embedding from the first assistant device and the second language speech recognition model when generating a corresponding recognition of the corresponding speech.
[0130] In some embodiments, the corresponding subset locally stored on the first assistant device includes a speech recognition model used to perform a first part of speech recognition. In some of these embodiments, assigning corresponding processing roles includes: assigning a first part of a language speech recognition role to the first assistant device, the first part of the language speech recognition role utilizing the speech recognition model when generating an output and transmitting the corresponding output to the second assistant device; and assigning a second language speech recognition role to the second assistant device, the second language speech recognition role performing beam search on the corresponding output from the first assistant device when generating a corresponding recognition of the corresponding speech.
[0131] In some embodiments, the corresponding subset locally stored on the first assistant device includes one or more pre-adaptation natural language understanding models that are used to perform semantic analysis of natural language inputs, and the one or more initial natural language understanding models occupy a first amount of local disk space at the first assistant device. In some of these embodiments, the corresponding subset locally stored on the first assistant device includes one or more post-adaptation natural language understanding models, the one or more post-adaptation natural language understanding models including at least one additional natural language understanding model in addition to the one or more initial natural language understanding models, and occupying a second amount of local disk space at the first assistant device, the second amount being greater than the first amount.
[0132] In some embodiments, the corresponding subset locally stored on the first assistant device includes a first device natural language understanding model that is used for semantic analysis for one or more first classifications and does not include any natural language understanding models that are used for semantic analysis for a second classification. In some of these embodiments, the corresponding subset locally stored on the second assistant device includes a second device natural language understanding model that is used for semantic analysis for at least the second classification.
[0133] In some embodiments, the corresponding processing capabilities of each assistant device among different assistant devices in the assistant device group include a corresponding processor value based on the capabilities of the processors on one or more devices, a corresponding memory value based on the size of the memory on the device, and a corresponding disk space value based on the available disk space.
[0134] In some embodiments, generating an assistant device group of different assistant devices is in response to a user interface input that explicitly indicates the desire to group different assistant devices.
[0135] In some embodiments, generating an assistant device group of different assistant devices is automatically performed in response to determining that the different assistant devices meet one or more proximity conditions relative to each other.
[0136] In some embodiments, generating an assistant device group of different assistant devices is performed in response to an affirmative user interface input that is received in response to a suggestion to create an assistant device group, and the suggestion is automatically generated in response to determining that the different assistant devices meet one or more proximity conditions relative to each other.
[0137] In some embodiments, the method further comprises: after assigning corresponding processing roles to each of the different assistant devices in the assistant device group, determining that a first assistant device is no longer in the group; and in response to determining that the first assistant device is no longer in the group, causing the first assistant device to replace a corresponding subset locally stored on the first assistant device with a first device-on model in a first set.
[0138] In some embodiments, determining the overall set is also based on usage data that reflects past usage at one or more of the assistant devices in the group.
[0139] In some of these embodiments, determining the overall set includes: determining a plurality of candidate sets, each of which can be collectively locally stored and collectively locally used by the assistant devices in the group, based on the corresponding processing capabilities of each of the different assistant devices in the assistant device group; and selecting the overall set from the candidate sets based on the usage data.
[0140] In some embodiments, a method implemented by one or more processors of an assistant device is provided. The method includes: in response to determining that the assistant device is included in an assistant device group that includes the assistant device and one or more additional assistant devices, determining that the assistant device is the leading device in the group. The method further includes: in response to determining that the assistant device is the leading device in the group, determining an overall set of device-on models for collaboratively locally processing assistant requests directed to any one of the different assistant devices in the assistant device group, based on the processing capabilities of the assistant device and based on the received processing capabilities of each of the one or more additional assistant devices; and for each device-on model in the overall set of device-on models, determining a corresponding assignment as to which of the assistant devices in the group will locally store the device-on model. The method further includes: in response to determining that the assistant device is the leading device in the group: communicating with the one or more additional assistant devices to cause each of the one or more additional assistant devices to locally store any one of the device-on models having the corresponding assignment for the additional assistant device; locally storing at the assistant device the device-on model having the corresponding assignment for the assistant device; and assigning one or more corresponding processing roles to each of the assistant devices in the group for collaboratively locally processing assistant requests directed to the group.
[0141] These and other embodiments of the techniques disclosed herein may optionally include one or more of the following features.
[0142] In some embodiments, determining that the assistant device is the leading device in the group includes: comparing the processing capabilities of the assistant device with the received processing capabilities of each of one or more additional assistant devices; and determining, based on the comparison, that the assistant device is the leading device in the group.
[0143] In some embodiments, the group of assistant devices is created in response to a user interface input that explicitly indicates a desire to group different assistant devices.
[0144] In some embodiments, the method further includes: in response to determining that the assistant device is the leading device in the group, and in response to receiving an assistant request at one or more of the assistant devices in the group, coordinating collaborative local processing of the assistant request using the corresponding processing role assigned to the assistant device.
[0145] In some embodiments, a method implemented by one or more processors of an assistant device is provided and includes determining that the assistant device has been removed from a group of different assistant devices. The group is a group that has included the assistant device and at least one additional assistant device. When the assistant device is removed from the group, a set of on-device models is stored locally on the assistant device, and the set of on-device models is insufficient to fully process locally at the assistant device an oral discourse directed to the automated assistant. The method further includes: in response to determining that the assistant device has been removed from the group of assistant devices: causing the assistant device to purge one or more of the on-device models in the set, and retrieve and locally store one or more additional on-device models. After retrieving and locally storing the one or more additional on-device models of the assistant device, the one or more additional on-device models and any remaining on-device models in the set of on-device models can be used to fully process locally at the assistant device an oral discourse directed to the automated assistant.
Claims
1. A method implemented by one or more processors, the method comprising: generating a group of assistant devices of different assistant devices, the different assistant devices including at least a first assistant device and a second assistant device, wherein, when generating the group: the first assistant device includes a first set of device - on - board models utilized when locally processing assistant requests directed to the first assistant device, and the second assistant device includes a second set of device - on - board models utilized when locally processing assistant requests directed to the second assistant device; determining, based on the corresponding processing capabilities of each of the different assistant devices in the assistant device group, an overall set of stored device - on - board models for collaboratively processing assistant requests directed to any one of the different assistant devices in the assistant device group, wherein the overall set of stored device - on - board models includes a first subset of device - on - board models and a second subset of device - on - board models, wherein the first subset includes at least one device - on - board model not included in the second subset, and the second subset includes at least one device - on - board model not included in the first subset; in response to generating the assistant device group: causing each of the different assistant devices to locally store a corresponding subset of the overall set of stored device - on - board models, wherein the first subset is locally stored on the first assistant device and the second subset is locally stored on the second assistant device, wherein causing the first assistant device to locally store the first subset includes: causing the first assistant device to clear one or more first device - on - board models in the first set to provide storage space for the first subset locally stored on the first assistant device, and wherein causing the second assistant device to locally store the second subset includes: causing the second assistant device to clear one or more second device - on - board models in the second set to provide storage space for the second subset locally stored on the second assistant device, wherein the first subset locally stored on the first assistant device omits the one or more cleared first device - on - board models, wherein the second subset locally stored on the second assistant device omits the one or more cleared second device - on - board models, and wherein the first subset locally stored on the first assistant device includes a given device - on - board model that is more accurate and / or more robust than a given corresponding device - on - board model included in the one or more cleared first device - on - board models; assigning one or more corresponding processing roles to each of the different assistant devices in the assistant device group, each of the processing roles utilizing one or more corresponding locally stored device - on - board models, and a given one of the processing roles being exclusively assigned to the first assistant device; and After generating the assistant device group and after assigning the corresponding processing roles to each of the different assistant devices in the assistant device group: Detect oral discourse via a microphone of at least one of the different assistant devices in the assistant device group, and In response to the oral discourse, cause the oral discourse to be collaboratively processed by the different assistant devices in the assistant device group using their corresponding processing roles.
2. The method according to claim 1, wherein causing the first assistant device to clear one or more first device - on models in the first set includes: causing the first assistant device to clear the first device wake - word detection model in the first set that is used to detect the first wake - word; wherein the second subset locally stored on the second assistant device includes a second device wake - word detection model that is used to detect the first wake - word; wherein assigning the corresponding processing role includes: assigning a first wake - word detection role to the second assistant device, the first wake - word detection role using the second device wake - word detection model when monitoring for the occurrence of the first wake - word; wherein the oral discourse includes a first wake - word followed by an assistant command; and wherein, in the first wake - word detection role, the second assistant device detects the occurrence of the first wake - word and causes an additional processing role in the corresponding processing role to be executed in response to detecting the occurrence of the first wake - word.
3. The method according to claim 2, wherein the additional processing role in the corresponding processing role is executed by the first assistant device, and wherein the second assistant device causes the additional processing role in the corresponding processing role to be executed by transmitting an indication of the detection of the first wake - word to the first assistant device.
4. The method according to claim 1, wherein the first subset locally stored on the first assistant device includes a first device first wake - word detection model that is used to detect one or more first wake - words and does not include any wake - word detection model that is used to detect one or more second wake - words; wherein the second subset locally stored on the second assistant device includes a second device second wake - word detection model that is used to detect one or more second hot - words and does not include any wake - word detection model that is used to detect the one or more first wake - words; wherein assigning the corresponding processing role includes: assigning a first wake - word detection role to the first assistant device, the first wake - word detection role using the first device wake - word detection model when monitoring for the occurrence of the one or more first wake - words; and wherein assigning the corresponding processing role includes: assigning a second wake - word detection role to the second assistant device, the second wake - word detection role using the second device wake - word detection model when monitoring for the occurrence of the one or more second wake - words.
5. The method according to claim 1, wherein The first subset locally stored on the first assistant device includes a first language speech recognition model for performing speech recognition of a first language and does not include any speech recognition model for recognizing speech of a second language; wherein, the second subset locally stored on the second assistant device includes a second language speech recognition model for performing speech recognition of the second language and does not include any speech recognition model for recognizing speech of the first language; wherein, assigning the corresponding processing role includes: assigning a first language speech recognition role to the first assistant device, the first language speech recognition role utilizing the first language speech recognition model when performing speech recognition of the first language; and wherein, assigning the corresponding processing role includes: assigning a second language speech recognition role to the second assistant device, the second language speech recognition role utilizing the second language speech recognition model when performing speech recognition of the second language.
6. The method according to claim 1, wherein, the first subset locally stored on the first assistant device includes a first part of a speech recognition model for performing a first part of speech recognition and does not include a second part of the speech recognition model; wherein, the second subset locally stored on the second assistant device includes the second part of the speech recognition model for performing the second part of speech recognition and does not include the first part of the speech recognition model; wherein, assigning the corresponding processing role includes: assigning a first part of a language speech recognition role to the first assistant device, the first part of the language speech recognition role utilizing the first part of the speech recognition model when generating a corresponding embedding of the corresponding speech and transmitting the corresponding embedding to the second assistant device; and wherein, assigning the corresponding processing role includes: assigning a second language speech recognition role to the second assistant device, the second language speech recognition role utilizing the corresponding embedding from the first assistant device and the second language speech recognition model when generating a corresponding recognition of the corresponding speech.
7. The method according to claim 1, wherein, the first subset locally stored on the first assistant device includes a speech recognition model for performing a first part of speech recognition; wherein, assigning the corresponding processing role includes: assigning a first part of a language speech recognition role to the first assistant device, the first part of the language speech recognition role utilizing the speech recognition model when generating an output and transmitting a corresponding output to the second assistant device; and wherein, assigning the corresponding processing role includes: assigning a second language speech recognition role to the second assistant device, the second language speech recognition role performing a beam search on the corresponding output from the first assistant device when generating a corresponding recognition of the corresponding speech.
8. The method according to claim 1, wherein, The first subset locally stored on the first assistant device includes one or more pre-adaptation natural language understanding models that are used to perform semantic analysis of natural language input, wherein the one or more initial natural language understanding models occupy a first amount of local disk space at the first assistant device; and wherein the first subset locally stored on the first assistant device includes one or more post-adaptation natural language understanding models, and the one or more post-adaptation natural language understanding models include at least one additional natural language understanding model in addition to the one or more initial natural language understanding models, wherein the one or more post-adaptation natural language understanding models occupy a second amount of local disk space at the first assistant device, and the second amount is greater than the first amount.
9. The method according to claim 1, wherein, the first subset locally stored on the first assistant device includes a first device natural language understanding model that is used for semantic analysis of one or more first classifications, and does not include any natural language understanding model that is used for semantic analysis of a second classification; and wherein the second subset locally stored on the second assistant device includes a second device natural language understanding model that is used for semantic analysis of at least the second classification.
10. The method according to claim 1, wherein, the corresponding processing capabilities of each of the different assistant devices in the group of assistant devices include a corresponding processor value based on the capabilities of the processors on one or more devices, a corresponding memory value based on the size of the memory on the device, and a corresponding disk space value based on the available disk space.
11. The method according to any one of claims 1 to 10, wherein, the group of assistant devices of different assistant devices is generated in response to a user interface input that explicitly indicates the expectation of grouping the different assistant devices.
12. The method according to any one of claims 1 to 10, wherein, the group of assistant devices of different assistant devices is automatically executed in response to determining that the different assistant devices meet one or more proximity conditions relative to each other.
13. The method according to any one of claims 1 to 10, further comprising: after assigning the corresponding processing roles to each of the different assistant devices in the group of assistant devices: determining that the first assistant device is no longer in the group; and in response to determining that the first assistant device is no longer in the group: causing the first assistant device to replace the first subset locally stored on the first assistant device with the first device model in the first set.
14. The method according to any one of claims 1 to 10, wherein, determining the overall set is also based on usage data that reflects past usage at one or more of the assistant devices in the group.
15. The method according to claim 14, wherein, determining the overall set includes: Determine a plurality of candidate sets that can be collectively locally stored and collectively locally used by the assistant devices in the group, based on the corresponding processing capabilities of each of the different assistant devices in the assistant device group; and Select the overall set from the candidate sets based on the usage data.
16. The method according to claim 1, wherein before generating the group, the first assistant device is capable of processing a given oral assistant request by itself, and wherein after generating the assistant device group and after assigning the corresponding processing roles to each of the different assistant devices in the assistant device group, the first assistant device is unable to process the given oral assistant request by itself.
17. A system for dynamically adapting an on-device model, comprising: A first assistant device; A second assistant device; A memory storing instructions; One or more processors operable to execute the instructions to: Generate a group of assistant devices of different assistant devices, the different assistant devices including at least the first assistant device and the second assistant device, wherein when generating the group: The first assistant device includes a first set of on-device models stored locally for use when locally processing assistant requests directed to the first assistant device, and The second assistant device includes a second set of on-device models stored locally for use when locally processing assistant requests directed to the second assistant device; Based on the corresponding processing capabilities of each of the different assistant devices in the assistant device group, determine an overall set of stored on-device models for collaboratively processing assistant requests directed to any one of the different assistant devices in the assistant device group, wherein the overall set of stored on-device models includes a first subset of on-device models and a second subset of on-device models, wherein the first subset includes at least one on-device model not included in the second subset, and the second subset includes at least one on-device model not included in the first subset; In response to generating the assistant device group: Cause each of the different assistant devices to locally store a corresponding subset of the overall set of stored on-device models, wherein the first subset is locally stored on the first assistant device and the second subset is locally stored on the second assistant device, wherein causing the first assistant device to locally store the first subset includes: causing the first assistant device to clear one or more first on-device models in the first set to provide storage space for the first subset locally stored on the first assistant device, and wherein causing the second assistant device to locally store the second subset includes: causing the second assistant device to clear one or more second on-device models in the second set to provide storage space for the second subset locally stored on the second assistant device, Among them, the first subset locally stored on the first assistant device omits the one or more on-device models on the cleared first devices. Among them, the second subset locally stored on the second assistant device omits the one or more on-device models on the cleared second devices, and Among them, the first subset locally stored on the first assistant device includes an on-device model for a given device, and the on-device model for the given device is more accurate and / or more robust than the corresponding on-device model for the given device included in the one or more on-device models on the cleared first devices. Assign one or more corresponding processing roles to each of the different assistant devices in the assistant device group, each of the processing roles utilizes one or more corresponding locally stored on-device models among the locally stored on-device models, and a given one of the processing roles is exclusively assigned to the first assistant device; and After generating the assistant device group and after assigning the corresponding processing roles to each of the different assistant devices in the assistant device group: Detect dictated speech via a microphone of at least one of the different assistant devices in the assistant device group, and In response to the dictated speech, cause the dictated speech to be collaboratively processed by the different assistant devices in the assistant device group using their corresponding processing roles.
18. The system according to claim 17, wherein, Before generating the group, the first assistant device is capable of processing a given dictated assistant request by itself, and wherein after generating the assistant device group and after assigning the corresponding processing roles to each of the different assistant devices in the assistant device group, the first assistant device is unable to process the given dictated assistant request by itself.
19. The system according to claim 17, wherein, The first assistant device includes the memory and the one or more processors.
Citation Information
Patent Citations
Multiple pass automatic speech recognition methods and apparatus
US20150058018A1
System and method for distributed virtual assistant platforms
US20150215350A1
Virtual assistant configured by selection of wake-up phrase
US20180108343A1
Collaborative voice controlled devices
US20180182397A1