Dynamically adapting on-device models, of grouped assistant devices, for cooperative processing of assistant requests
By adapting on-device models and processing roles among a group of assistant devices based on individual capabilities, the solution addresses processing limitations, enhancing robustness and accuracy while reducing latency and network usage.
Patent Information
- Application Number
- JP2025065107
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2020-11-13
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-30
- Estimated Expiration
- 2040-12-11
AI Technical Summary
Existing assistant devices face limitations in processing power and memory capacity, leading to less robust and accurate local components, particularly in older and less expensive devices, necessitating reliance on cloud-based counterparts for processing, which increases latency and network usage.
Dynamically adapt on-device models and processing roles among a group of assistant devices, considering individual capabilities, enabling cooperative processing to enhance robustness and accuracy, reducing the need for remote data transmission and improving security.
The distributed pipeline of assistant components achieves increased robustness and accuracy, reducing latency and network usage, enhancing data security by processing requests locally with improved efficiency.
Smart Images

Figure 2025111525000001_ABST
Abstract
Description
Background Art
[0001] Humans can engage in human-computer dialogues with interactive software applications referred to herein as "automation assistants" (also referred to as "chatbots", "interactive personal assistants", "intelligent personal assistants", "personal voice assistants", "conversational agents", etc.). For example, a human (who may be referred to as a "user" when interactively operating an automation assistant) can provide commands and / or requests to an automation assistant using natural language voice input (i.e., a spoken utterance), which in some cases can be converted to text and then processed. The commands and / or requests can, in addition or alternatively, be provided via one or more other input modalities such as text (e.g., typed) natural language input, touch screen input, and / or touch-free gesture input (e.g., detected by a camera of a corresponding assistant device). An automation assistant generally responds to a command or request by providing response user interface output (e.g., auditory and / or visual user interface output), controlling a smart device, and / or performing other actions.
[0002] Automated assistants typically rely on a pipeline of components when processing user requests. For example, a wake word detection engine can process audio data when monitoring for the occurrence of an audio wake word (e.g., "OK Assistant") and cause other components to perform processing in response to detection of the occurrence. As another example, an automatic speech recognition (ASR) engine can be used to process audio data that includes spoken utterances to generate a transcript of the user's utterance (i.e., a sequence of words and / or other tokens). The ASR engine can process the audio data based on the next occurrence of an audio wake word when detected by the wake word detection engine and / or in response to other invocations of the automated assistant. As another example, a natural language understanding (NLU) engine can be used to process the text of a request (e.g., text converted from a spoken utterance using ASR) to generate a symbolic representation, or belief state, that is a semantic representation of the text. For example, the belief state can include an intent corresponding to the text and optionally parameters (e.g., slot values) for the intent. The belief state represents the actions to be performed in response to the spoken utterance after it has been fully formed over one or more dialog turns (e.g., after all required parameters have been resolved). Another fulfillment component can then utilize the fully formed belief state to execute the actions corresponding to the belief state.
[0003] When interacting with an automated assistant, a user utilizes one or more assistant devices (client devices having an automated assistant interface). The pipeline of components utilized when processing requests provided at the assistant device can include components executed locally at the assistant device and / or components executed at one or more remote servers that are network communicating with the assistant device.
[0004] Efforts have been made to increase the number of components executed locally on the assistant device and / or to enhance the robustness and / or accuracy of such components. Considerations such as reducing latency, improving data security, reducing network usage, and / or achieving other technical advantages motivate those efforts. As an example, some assistant devices can include a local wake word engine and / or a local ASR engine.
[0005] However, due to the limited processing capabilities of various assistant devices, components implemented locally on the assistant device may be less robust and / or accurate than their cloud-based counterparts. This can apply particularly to older and / or less expensive assistant devices that may lack (a) the processing power and / or memory capacity to execute various components and / or utilize their associated models, and (b) and / or the disk space available to store various associated models. SUMMARY OF THE INVENTION PROBLEM TO BE SOLVED BY THE INVENTION
[0006] The implementations disclosed herein are directed to dynamically adapting an assistant on-device model stored locally on an assistant device of an assistant device group and / or adapting an assistant processing role of an assistant device of an assistant device group. In some of those implementations, the corresponding on-device model and / or the corresponding processing role are determined based on collectively considering the individual processing capabilities of the assistant devices of the group for each assistant device of the group. For example, the on-device model and / or the processing role for a given assistant device may be determined considering the corresponding processing capabilities of other assistant devices of the group (e.g., other devices can store other required on-device models and / or execute other required processing roles) based on the individual processing capabilities of the given assistant device (e.g., considering the constraints of the processor, memory, and / or storage, the given assistant device can store those on-device models and execute those processing roles). In some implementations, usage data can also be utilized in determining the corresponding on-device model and / or the corresponding processing role for each of the assistant devices of the group.
[0007] The implementations disclosed herein are also, alternatively, directed to cooperatively utilizing the assistant devices of the group, as well as their associated post-adaptation on-device models and / or post-adaptation processing roles, when cooperatively processing an assistant request directed to any one of the assistant devices of the group.
Means for Solving the Problems
[0008] In these and other ways, the on-device models and on-device processing roles can be distributed among a plurality of heterogeneous assistant devices in a group, taking into account the processing capabilities of those assistant devices. Further, the collective robustness and / or capabilities of the on-device models and on-device processing roles exceed what is individually possible by any one of the assistant devices when distributed among a plurality of assistant devices in a group. Said another way, the implementations disclosed herein can effectively implement an on-device pipeline of assistant components that are distributed among a group of assistant devices. The robustness and / or accuracy of such a distributed pipeline well exceeds the robustness and / or accuracy capabilities of any pipeline implemented on only a single one of the group of assistant devices. By increasing robustness and / or accuracy as disclosed herein, the latency for a greater number of assistant requests can be reduced. Further, as a result of increased robustness and / or accuracy, less (or no) data is transmitted to remote automated assistant components when resolving assistant requests. As a direct result of this, the security of user data is improved, the frequency of network use is reduced, the amount of data transmitted over the network is reduced, and / or the latency when resolving assistant requests is reduced (e.g., those assistant requests that are resolved locally can be resolved with less latency than when remote assistant components are involved).
[0009] As a result of some adaptations, one or more assistant devices in a group may lack the engines and / or models necessary for that assistant device to process many assistant requests alone. For example, as a result of some adaptations, an assistant device may lack any wake word detection ability and / or ASR ability. However, when within a group and adapted according to the implementations disclosed herein, an assistant device can cooperate with other assistant devices in the group to each perform its own processing role and, in doing so, utilize its own on-device model to cooperatively process assistant requests. Thus, an assistant request directed to an assistant device can be processed as is, in cooperation with other assistant devices in the group.
[0010] In various implementations, adaptation of the group to the assistant devices is performed in response to group generation or group modification (e.g., incorporation or removal of assistant devices from the group). As described herein, groups can be generated based on explicit user input indicating a desire for the group and / or can be automatically generated, for example, based on determining that the assistant devices of the group meet proximity conditions with respect to each other. In implementations that perform adaptation only when a group is created in such a manner, the occurrence of two assistant requests received simultaneously at two separate devices of the group (and which may not be able to be processed cooperatively in parallel) can be mitigated. For example, when proximity conditions are considered in generating a group, it is unlikely that two different simultaneous requests will be received at two different assistant devices of the group. For example, this is less likely to occur when the assistant devices of the group are all in the same room as opposed to being spread across multiple floors of a house. As another example, when user input explicitly indicates that a group should be created, non-overlapping assistant requests are likely to be provided to the assistant devices of the group.
[0011] As a specific example of various implementation forms, assume that an assistant device group consisting of a first assistant device and a second assistant device is generated. Further, at the time when the assistant device group is generated, assume that the first assistant device includes a wake word engine and a corresponding wake word model, a warm cue engine and a corresponding warm cue model, an authentication engine and a corresponding authentication model, as well as a corresponding on-device ASR engine and a corresponding ASR model. Further, at the time when the assistant device group is generated, assume that the second assistant device also includes the same engines and models as the first assistant device (or a modified form thereof), and in addition, includes an on-device NLU engine and a corresponding NLU model, an on-device fulfillment and a corresponding fulfillment model, as well as an on-device TTS engine and a corresponding TTS model.
[0012] In response to the first assistant device and the second assistant device being grouped, the assistant on-device models stored locally in the first assistant device and the second assistant device, and / or the corresponding processing roles of the first and second assistant devices may be adapted. The on-device models and processing roles may be determined based on considering the first processing capabilities of the first assistant device and the first processing capabilities of the second assistant device for each of the first and second assistant devices. For example, a set of on-device models including a first subset that can be stored and utilized on the first assistant device may be determined by the corresponding processing role / engine. Further, the set may be able to include a second subset that can be stored and utilized on the second assistant device by the corresponding processing role. For example, the first subset can include only an ASR model, but the ASR model of the first subset can be more robust and / or accurate than the pre-adaptation ASR model of the first assistant device. Further, they may need to utilize more extensive computing resources when performing ASR. However, the first assistant device can have only the ASR model and ASR engine of the first subset, and the pre-adaptation model and engine can be purged, thereby freeing up the computing resources utilized when performing ASR using the ASR model of the first subset. Continuing with this example, the second subset can include the same models as the second assistant device previously included, except that the ASR model can be omitted and a more robust and / or accurate NLU model can replace the pre-adaptation NLU model. The more robust and / or accurate NLU models may require more resources compared to the pre-adaptation NLU models, but they can be freed up through the purging of the pre-adaptation ASR model (and omitting any ASR models from the second subset).
[0013] Next, when the first assistant device and the second assistant device cooperate to process an assistant request directed to any of the group of assistant devices, they can cooperatively utilize their associated post - adaptation on - device models and / or post - adaptation processing roles. For example, assume there is a spoken utterance of "OK Assistant, increase the temperature two degrees". The wake cue engine of the second assistant device can detect the occurrence of the wake cue "OK Assistant". In response, the wake cue engine can transmit a command to the ASR engine of the first assistant device to perform speech recognition on the captured audio data following the wake cue. The transcript generated by the ASR engine of the first assistant device may be transmitted to the second assistant device, and the NLU engine of the second assistant device can perform NLU on the transcript. The results of the NLU may be communicated to the fulfillment engine of the second assistant device, which can use those NLU results to determine the command to increase the temperature by two degrees to be transmitted to the smart thermostat.
[0014] The foregoing content has been presented as an overview of only some implementations. These and other implementations are disclosed in more detail herein.
[0015] Furthermore, some implementations may include a system having one or more user devices, each device including one or more processors and a memory operably coupled to the one or more processors. The memory of the one or more user devices may store instructions that, in response to execution of instructions by the one or more processors of the one or more user devices, cause the one or more processors to execute any of the methods described herein. Some implementations may also include at least one non-transitory computer-readable medium including instructions that, in response to execution by one or more processors, cause the one or more processors to execute any of the methods described herein.
[0016] It should be understood that all combinations of the foregoing concepts and additional concepts described in greater detail herein are intended to be part of the subject matter disclosed herein. For example, all combinations of claimed subject matter appearing at the end of this disclosure are intended to be part of the subject matter disclosed herein.
Brief Description of the Drawings
[0017]
Figure 1A
Figure 1B1
Figure 1B2
Figure 1B3
Figure 1C
Figure 1D
Figure 2
Figure 3
Figure 4
[0018] Many users may interact with an automated assistant using any one of a plurality of assistant devices. For example, some users may have, among other assistant devices, one or more smartphones, one or more tablet computers, one or more vehicle computing systems, one or more wearable computing devices, one or more smart TVs, one or more interactive stand-alone speakers, one or more interactive stand-alone speakers with a display, one or more IoT devices, etc., that can receive user input directed to the automated assistant and / or can be controlled by the automated assistant, and may own a collaborative "ecosystem" of assistant devices.
[0019] A user can engage in a human-computer dialog with the automated assistant using any of these assistant devices (assuming that an automated assistant client is installed and the assistant device can receive input). In some cases, these assistant devices may be scattered around the user's primary residence, vacation home, workplace, and / or other structures. For example, mobile assistant devices such as smartphones, tablets, smartwatches, etc., may be worn by the user and / or may be in the location where the user last placed them. Other assistant devices such as conventional desktop computers, smart TVs, interactive stand-alone speakers, and IoT devices may be relatively stationary, but nevertheless may be placed in various locations (e.g., rooms) within the user's home or workplace.
[0020] Referring initially to FIG. 1A, an exemplary assistant ecosystem is illustrated. The exemplary assistant ecosystem includes a first assistant device 110A, a second assistant device 110B, a third assistant device 110C, and a fourth assistant device 110D. The assistant devices 110A-D can all be disposed within a home, enterprise, or other environment. Further, the assistant devices 110A-D can all be linked together in one or more data structures or otherwise associated with each other. For example, the four assistant devices 110A-D can all be registered with the same user account, registered with the same set of user accounts, registered in a particular structure, and / or all assigned to a particular structure in a device topology representation. The device topology representation can include a corresponding unique identifier for each of the assistant devices 110A-D and optionally can include a corresponding unique identifier for other devices that are not assistant devices (but can be interactively operated via an assistant device), such as, for example, an IoT device that does not include an assistant interface. Further, the device topology representation can specify device attributes associated with each of the assistant devices 110A-D. The device attributes for a given assistant device can indicate, for example, one or more input and / or output modalities supported by each of the assistant devices, the processing capabilities of each of the assistant devices, the manufacturing form factor, model, and / or unique identifier (e.g., serial number) of each of the assistant devices (based on which the processing capabilities can be determined), and / or other attributes. As another example, the four assistant devices can all be linked together or otherwise associated with each other as a function of being connected to the same wireless network, such as a secure access wireless network, and / or as a function of communicating peer-to-peer with each other collectively (e.g., via Bluetooth after pairing).In other words, in some implementations, multiple assistant devices are linked together as a function of making a secure network connection with each other and may potentially be adapted by the implementations disclosed herein, without necessarily being associated with each other in any data structure.
[0021] As a non-limiting example, the first assistant device 110A may be a first type of assistant device, such as a particular model of an interactive stand-alone speaker with a display and a camera. The second assistant device 110B may be a second type of assistant device, such as a first model of an interactive stand-alone speaker without a display or a camera. The assistant devices 110C and 110D may each be a third type of assistant device, such as a third model of an interactive stand-alone speaker without a display. The third type (assistant devices 110C and 110D) may be inferior in processing power compared to the second type (assistant device 110D). For example, the third type may have a processor that is inferior in processing power compared to the processor of the second type. For example, the processor of the third type may not include a GPU, while the processor of the first type includes a GPU. Also, for example, the processor of the third type may have a smaller cache and / or a lower operating frequency compared to the processor of the second type. As another example, the size of the on-device memory of the third type can be made smaller than the size of the on-device memory of the second type (e.g., 1GB compared to 2GB). As yet another example, the available disk space of the third type can be made smaller than the available disk space of the first type. The available disk space may be different from the currently available disk space. For example, the available disk space can be determined as the currently available disk space plus the disk space currently occupied by one or more on-device models. As another example, the available disk space can be the total disk space minus the space occupied by the operating system and / or other specific software. Continuing with this example, the first type and the second type can have the same processing power.
[0022] In addition to being linked together in a data structure, two or more (e.g., all) of the assistant devices 110A - D communicate with each other at least selectively via a local area network (LAN) 108. The LAN 108 can include a wireless network such as one that utilizes Wi-Fi, a direct peer-to-peer network such as one that utilizes Bluetooth, and / or other communication topologies that utilize other communication protocols.
[0023] Assistant device 110A includes an assistant client 120A, which can be a stand-alone application on top of an operating system or can form all or part of the operating system of assistant device 110A. Assistant client 120A includes, in FIG. 1A, a wake / call engine 121A1 and one or more associated on-device wake / call models 131A1. Wake / call engine 121A1 monitors the occurrence of one or more wake or call queues and, in response to detecting one or more queues, can call one or more previously inactive functions of assistant client 120A. For example, calling assistant client 120A can include activating an ASR engine 122A1, an NLU engine 123A1, and / or other engines. For example, this can cause the ASR engine 122A1 to process additional audio data frames following a wake or call queue (whereas prior to the call, no further processing of audio data frames occurred), and / or cause assistant client 120A to transmit additional audio data frames and / or other data to be transmitted for processing to a cloud-based assistant component 140 (e.g., processing of audio data frames by a remote ASR engine of the cloud-based assistant component 140).
[0024] In some implementations, the wake queue engine 121A continuously processes a stream of audio data frames based on the output from one or more microphones of the client device 110A (e.g., when not in the "inactive" mode) and can monitor for the occurrence of an audio wake word or call phrase (e.g., "OK Assistant", "Hey Assistant"). The processing can be performed by the wake queue engine 121A utilizing one or more of the wake models 131A1. For example, one of the wake models 131A1 can be a neural network model trained to process frames of audio data and generate an output indicating whether one or more wake words are present in the audio data. While monitoring for the occurrence of a wake word, the wake queue engine 121 discards any audio data frames that do not contain the wake word (e.g., after temporary storage in a buffer). In addition to, or instead of, monitoring for the occurrence of a wake word, the wake queue engine 121A1 can monitor for the occurrence of other call queues. For example, the wake queue engine 121A1 can also monitor for the pressing of a call hardware button and / or a call software button. As another example, in a continuing embodiment, when the assistant device 110A includes a camera, the wake queue engine 121A1 can optionally process image frames from the camera when monitoring for the occurrence of a call gesture such as a wave while the user's gaze is directed at the camera and / or the occurrence of other call queues such as the user's gaze being directed at the camera along with an indication that the user is speaking.
[0025] Assistant client 120A also includes, in FIG. 1A, an automatic speech recognition (ASR) engine 122A1 and one or more associated on-device ASR models 132A1 in FIG. 1A. The ASR engine 122A1 processes speech data including uttered speech and can be used to generate a transcript of the user's speech (i.e., a sequence of words and / or other tokens). The ASR engine 122A1 can process the speech data using the on-device ASR model 132A1. The on-device ASR model 132A1 can include, for example, a two-pass ASR model, which is a neural network model and is utilized by the ASR engine 122A1 to generate a sequence of probabilities (and the probabilities utilized to generate a transcript) over tokens. As another example, the on-device ASR model 132A1 can include an acoustic model that is a neural network model and a language model that includes a mapping of a sequence of phonemes to words. The ASR engine 122A1 can process the speech data using the acoustic model to generate a sequence of phonemes and map the sequence of phonemes to specific words using the language model. Additional or alternative ASR models may be utilized.
[0026] Assistant client 120A also includes, in FIG. 1A, a natural language understanding (NLU) engine 123A1 and one or more associated on-device NLU models 133A1 in FIG. 1A. The NLU engine 123A1 can generate a symbolic representation, which is a semantic representation of natural language text such as the text of a transcript generated by the ASR engine 122A1 or typed input text (e.g., typed using the virtual keyboard of the assistant device 110A), or a belief state. For example, the belief state can include an intent corresponding to the text and optionally parameters (e.g., slot values) for the intent. The belief state represents an action to be performed in response to an utterance after being fully formed through one or more dialog turns (e.g., after all required parameters are resolved). When generating the symbolic representation, the NLU engine 123A1 can utilize one or more on-device NLU models 133A1. The NLU model 133A1 can include one or more neural network models trained to process text and generate an output indicating the intent represented by the text and / or an indication of which part of the text corresponds to which parameter of the intent. The NLU model can additionally or alternatively include one or more models including a mapping to the corresponding symbolic representation of the text and / or template. For example, these mappings can include a mapping of the text "what time is it" to the intent of "current time" with the parameter of "current location". As another example, the mapping can include a mapping of the template "add [item(s)] to my shopping list" to the intent of "insert in shopping list" with the parameter of the item(s) included in the actual natural language corresponding to [item(s)] in the template.
[0027] In Figure 1A, the Assistant Client 120A also includes a Fulfillment Engine 124A1 and one or more associated on-device fulfillment models 134A1. The Fulfillment Engine 124A1 can use the fully formed symbolic representation to execute or cause to execute an action corresponding to the symbolic representation from the NLU Engine 123A1. The action can include providing a responsive user interface output (e.g., auditory and / or visual user interface output), controlling a smart device, and / or performing other actions. When executing or causing to execute an action, the Fulfillment Engine 124A1 can utilize the fulfillment model 134A1. As an example, for an intent of "turn on" with parameters specifying a particular smart light, the Fulfillment Engine 124A1 can use the fulfillment model 134A1 to identify the network address of the particular smart light and / or the command to be transmitted to transition the particular smart light to the "on" state. As another example, for an intent of "current" with a parameter of "current location", the Fulfillment Engine 124A1 can use the fulfillment model 134A1 to identify that the current time on the client device 110A should be retrieved and (using the TTS Engine 125A1) rendered audibly.
[0028] Assistant client 120A also includes, in FIG. 1A, a text-to-speech (TTS) engine 125A1 and one or more associated on-device TTS models 135A1 in FIG. 1A. The TTS engine 125A1 can process text (or its audio representation) using the on-device TTS model 135A1 to generate synthesized speech. The synthesized speech can be aurally rendered via a speaker of the local text-to-speech (“TTS”) engine (which converts text to speech) of the assistant device 110A. The synthesized speech can be generated and rendered as all or part of a response from the automated assistant and / or when prompting the user to define and / or clarify parameters and / or intents (e.g., as orchestrated by the NLU engine 123A1 and / or another dialog state engine).
[0029] Assistant client 120A also includes, in FIG. 1A, an authentication engine 126A1 and one or more associated on-device authentication models 136A1. The authentication engine 126A1 can utilize one or more authentication techniques to verify which of a plurality of registered users is interactively operating the assistant device 110, or, if only a single user is registered to the assistant device 110, the registered user (or guest / unregistered user instead) who is interactively operating the assistant device 110. As an example, text-dependent speaker verification (TD-SV) can be generated and stored for each of the registered users (e.g., in relation to the corresponding user profile) with permission from the associated user. The authentication engine 126A1 can utilize the TD-SV model of the on-device authentication model 136A1 when generating the corresponding TD-SV and / or when processing the corresponding portion of the audio data TD-SV and subsequently comparing it with the stored TD-SV to determine if there is a match to generate the corresponding current TD-SV. As another example, the authentication engine 126A1 can alternatively or additionally utilize text-independent speaker verification (TI-SV) techniques, speaker verification techniques, face verification techniques, and / or other verification techniques (e.g., PIN entry), and the corresponding on-device authentication model 136A1 when authenticating a particular user.
[0030] In Figure 1A, the Assistant Client 120A also includes a Warm Queue Engine 127A1 and one or more associated on-device warm queue models 137A1. The Warm Queue Engine 127A1 at least selectively monitors for the occurrence of one or more warm words or other warm queues and, in response to detecting one or more of the warm queues, can cause a particular action to be performed by the Assistant Client 120A. A warm queue may be in addition to any wake word or other wake queue, and each of the warm queues can be at least selectively active. In particular, detecting the occurrence of a warm queue causes a particular action to be performed even when there is no wake queue prior to the detected occurrence. Thus, when a warm queue is a particular one or more words, the user can simply speak that word without the need to utter a wake queue and can cause the execution of a corresponding particular action.
[0031] As an example, the "stop" warm cue may be active at least when a timer or alarm is being audibly rendered on the assistant device 110A via the automation assistant 120A. For example, in such a case, the warm cue engine 127A continuously (or at least when the VAD engine 128A1 detects voice activity) processes a stream of audio data frames based on the output from one or more microphones of the client device 110A to monitor for the occurrence of "stop", "halt", or other limited set of specific warm words. The processing can be performed by the warm cue engine 127A that utilizes one of the warm cue models 137A1, such as a neural network model trained to process frames of audio data and generate an output indicating whether the vocal occurrence of "stop" is present in the audio data. In response to detecting the occurrence of "stop", the warm cue engine 127A may cause a command to stop the sounding timer or alarm to be executed. In such a case, the warm cue engine 127A can continuously (or at least when the presence sensor detects presence) process a stream of images from the camera of the assistant device 110A to monitor for the occurrence of a hand in the "stop" gesture. The processing can be performed by the warm cue engine 127A that utilizes one of the warm cue models 137A1, such as a neural network model trained to process frames of vision data and generate an output indicating whether a hand is present in the "stop" gesture. In response to detecting the occurrence of the "stop" gesture, the warm cue engine 127A can cause a command to stop the sounding timer or alarm to be executed.
[0032] As another example, the "volume up", "volume down", and "next" warm cues may be active at least when music is being aurally rendered on the assistant device 110A via the automated assistant 120A. For example, in such a case, the warm cue engine 127A can continuously process a stream of audio data frames based on the output from one or more microphones of the client device 110A. The processing can include monitoring for the occurrence of "volume up" using a first one of the warm cue models 137A1, monitoring for the occurrence of "volume down" using a second one of the warm cue models 137A1, and monitoring for the occurrence of "next" using a third one of the warm cue models 137A1. In response to detecting the occurrence of "volume up", the warm cue engine 127A can cause a command to increase the volume of the rendered music to be executed, in response to detecting the occurrence of "volume down", the warm cue engine can cause a command to decrease the volume of the music to be executed, and in response to detecting the occurrence of "volume down", the warm cue engine can cause a command to be executed that causes the next track to be rendered instead of the current music track.
[0033] Assistant client 120A also includes, in FIG. 1A, a voice activity detector (VAD) engine 128A1 and one or more associated on-device VAD models 138A1 in FIG. 1A. The VAD engine 128A1 at least selectively monitors for the occurrence of voice activity in the voice data and, in response to detecting an occurrence, can cause one or more functions to be performed by the assistant client 120A. For example, the VAD engine 128A1 can activate the wake queue engine 121A1 in response to detecting voice activity. As another example, the VAD engine 128A1 may be utilized in a continuous listening mode, whereby it monitors for the occurrence of voice activity in the voice data and, in response to detecting an occurrence, can activate the ASR engine 122A1. The VAD engine 128A1 can process the voice data using the VAD model 138A1 when determining whether voice activity is present in the voice data.
[0034] Specific engines and corresponding models have been described so far with respect to assistant client 120A. However, it should be noted that some engines may be omitted and / or additional engines may be included. Also, it should be noted that assistant client 120A can fully process many assistant requests, including many assistant requests that are provided as spoken utterances, through its various on-device engines and corresponding models. However, since client device 110A is relatively restricted with respect to processing power, there are still many assistant requests that cannot be fully processed locally on assistant device 110A. For example, NLU engine 123A1 and / or corresponding NLU model 133A1 may only cover a subset of all available intents and / or parameters available via the automated assistant. As another example, fulfillment engine 124A1 and / or corresponding fulfillment model may only cover a subset of available fulfillments. As yet another example, ASR engine 122A1 and corresponding ASR model 132A1 may not have sufficient robustness and / or accuracy to correctly transcribe various spoken utterances.
[0035] Taking these and other considerations into account, the cloud-based assistant component 140 can be at least selectively utilized as needed when performing at least a portion of the processing of the assistant requests received at the assistant device 110A. The cloud-based automated assistant component 140 can include an engine and / or model (and / or additional or alternative forms) that pairs with the engine of the assistant device 110A. However, since the cloud-based automated assistant component 140 can leverage the virtually infinite resources of the cloud, one or more cloud-based counterparts may be more robust and / or accurate than those of the assistant client 120A. As an example, in response to an utterance requesting the execution of an assistant action not supported by the local NLU engine 123A1 and / or the local fulfillment engine 124A1, the assistant client 120A can transmit the voice data for the utterance, and / or its transcript generated by the ASR engine 122A1, to the cloud-based automated assistant component 140. The cloud-based automated assistant component 140 (e.g., its NLU engine and / or fulfillment engine) can perform more robust processing of such data and enable the resolution and / or execution of the assistant action. The transmission of data to the cloud-based automated assistant component 140 occurs via one or more wide area networks (WANs) 109 such as the Internet or a private WAN.
[0036] The second assistant device 110B includes an assistant client 120B, which can be a stand-alone application on top of the operating system or can form all or part of the operating system of the assistant device 110B. Similar to the assistant client 120A, the assistant client 120B includes a wake / call engine 121B1 and one or more associated on-device wake / call models 131B1, an ASR engine 122B1 and one or more associated on-device ASR models 132B1, an NLU engine 123B1 and one or more associated on-device NLU models 133B1, a fulfillment engine 124B1 and one or more associated on-device fulfillment models 134B1, a TTS engine 125B1 and one or more associated on-device TTS models 135B1, an authentication engine 126B1 and one or more associated on-device authentication models 136B1, a warm queue engine 127B1 and one or more associated on-device warm queue models 137B1, as well as a VAD engine 128B1 and one or more associated on-device VAD models 138B1.
[0037] Some or all of the engines and / or models of the Assistant Client 120B may be the same as those of the Assistant Client 120A, and / or some or all of the engines and / or models may be different. For example, the Wake Queue Engine 121B1 may not have the function to detect the wake queue in the image, and / or the Wake Model 131B1 may not have the model to process the image to detect the wake queue, while the Wake Queue Engine 121A1 has such a function and the Wake Model 131B1 includes such a model. This may be due to, for example, the Assistant Device 110A including a camera and the Assistant Device 110B not including a camera. As another example, the ASR model 131B1 used by the ASR engine 122B1 may be different from the ASR model 131A1 used by the ASR engine 122A1. This may be due to, for example, different models being optimized for different processor and / or memory capabilities between the Assistant Device 110A and the Assistant Device 110B.
[0038] Certain engines and corresponding models have been described so far with respect to the Assistant Client 120B. However, it should be noted that some engines may be omitted and / or additional engines may be included. Also, note that the Assistant Client 120B can fully process many assistant requests, including many assistant requests provided as voice utterances, through its various on-device engines and corresponding models. However, since the client device 110B is relatively restricted in terms of processing power, there are still many assistant requests that cannot be fully processed locally in the Assistant Device 110B. Considering these and other considerations, the cloud-based assistant component 140 may be at least selectively used when executing at least a part of the processing of the assistant requests received in the Assistant Device 110B.
[0039] The third assistant device 110C includes an assistant client 120C, which can be a stand-alone application on top of the operating system or form all or part of the operating system of the assistant device 110C. Similar to the assistant client 120A and the assistant client 120B, the assistant client 120C includes a wake / call engine 121C1 and one or more associated on-device wake / call models 131C1, an authentication engine 126C1 and one or more associated on-device authentication models 136C1, a warm queue engine 127C1 and one or more associated on-device warm queue models 137C1, and a VAD engine 128C1 and one or more associated on-device VAD models 138C1. Some or all of the engines and / or models of the assistant client 120C may be the same as those of the assistant client 120A and / or the assistant client 120B, and / or some or all of the engines and / or models may be different.
[0040] However, note that, unlike assistant clients 120A and 120B, assistant client 120C does not include any ASR engine or associated model, any NLU engine or associated model, any fulfillment engine or associated model, and any TTS engine or associated model. Further note that assistant client 120B can only fully process some assistant requests (i.e., those that match the warm queue detected by the warm queue engine 127C1) through its various on-device engines and corresponding models, and cannot process many assistant requests such as those provided as voiced utterances that do not match the warm queue. Considering these and other considerations, the cloud-based assistant component 140 can always be at least selectively utilized when executing at least partial processing of the assistant requests received at the assistant device 110C.
[0041] The fourth assistant device 110D includes an assistant client 120D, which can be a stand-alone application on top of the operating system or form all or part of the operating system of the assistant device 110D. Similar to assistant client 120A, assistant client 120B, and assistant client 120C, assistant client 120D includes a wake / call engine 121D1 and one or more associated on-device wake / call models 131D1, an authentication engine 126D1 and one or more associated on-device authentication models 136D1, a warm queue engine 127D1 and one or more associated on-device warm queue models 137D1, and a VAD engine 128D1 and one or more associated on-device VAD models 138D1. Some or all of the engines and / or models of assistant client 120C may be the same as those of assistant client 120A, assistant client 120B, and / or assistant client 120C, and / or some or all of the engines and / or models may be different.
[0042] However, note that, unlike assistant clients 120A and 120B, and similar to assistant client 120C, assistant client 120D does not include any arbitrary ASR engine or associated model, any arbitrary NLU engine or associated model, any arbitrary fulfillment engine or associated model, and any arbitrary TTS engine or associated model. Further note that assistant client 120D can only fully process some assistant requests (i.e., those that match the warm queue detected by warm queue engine 127D1) through its various on-device engines and corresponding models, and cannot process many assistant requests such as those provided as voice utterances and not matching the warm queue. Considering these and other considerations, the cloud-based assistant component 140 can always be at least selectively utilized when performing at least partial processing of the assistant requests received at assistant device 110D.
[0043] Next, referring to FIGS. 1B1, 1B2, 1B3, 1C, and 1D, different non-limiting examples of an assistant device group are illustrated along with different non-limiting examples of adaptations that may be implemented in response to the generation of the assistant device group. Through each of the adaptations, the grouped assistant devices may be well utilized collectively when processing various assistant requests, and through the collective utilization, they can execute the processing of those various assistant requests with higher robustness and / or accuracy compared to the case where any one of the assistant devices in the group could be executed individually before the adaptation. As a result, various technical advantages such as those described herein are obtained.
[0044] In FIGS. 1B1, 1B2, 1B3, 1C, and 1D, the engines and models of the assistant client having the same reference numbers as in FIG. 1A are not adapted with respect to FIG. 1A. For example, in FIGS. 1B1, 1B2, and 1B3, since the assistant client devices 110C and 110D are not included in group 101B of FIGS. 1B1, 1B2, and 1B3, the engines and models of the assistant client devices 110C and 110D are not adapted. However, in FIGS. 1B1, 1B2, 1B3, 1C, and 1D, the engines and models of the assistant client having reference numbers different from those in FIG. 1A (i.e., ending with "2", "3", or "4" instead of "1") indicate that it is adapted with respect to the corresponding object in FIG. 1. Further, an engine or model having a reference number ending with "2" in one figure and ending with "3" in the other figure means that different adaptations for the engine or model are made between the figures. Similarly, an engine or model having a reference number ending with "4" in the figure means that the adaptation for the engine or model is different from the case where the reference number ends with "2" or "3" in that figure.
[0045] Referring initially to FIG. 1B1, a device group 101B is created, and assistant devices 110A and 110B are included in device group 101B. In some implementations, device group 101B may be generated in response to a user interface input that explicitly indicates a desire to group assistant devices 110A and 110B. As an example, a user can provide a spoken utterance of "group [label for assistant device 110A] and [label for assistant device 110B]" to any one of assistant devices 110A - D. Such a spoken utterance is processed by the respective assistant device and / or a cloud - based assistant component 140 and interpreted as a request to group assistant devices 110A and 110B, and group 101B can be generated in response to such an interpretation. As another example, a registered user of assistant devices 110A - D can provide touch inputs in an application that enables configuration of the settings of assistant devices 110A - D. Those touch inputs can explicitly specify that assistant devices 110A and 110B should be grouped, and group 101B can be generated in response thereto. As yet another example, one of the exemplary techniques described below can be utilized instead when automatically generating device group 101B to determine that device group 101B should be generated, but a user input explicitly approving the generation of device group 101B can be required before generating device group 101B. For example, a prompt indicating that device group 101B should be generated can be rendered on one or more of assistant devices 110A - D, and device group 101B is actually generated only if (and optionally only if the user interface input is verified to come from a registered user) a positive user interface input is received in response to the prompt.
[0046] In some implementations, the device group 101B can instead be automatically generated. In some of those implementations, a user interface output indicating the generation of the device group 101B can be rendered by one or more of the assistant devices 110A - D to notify the corresponding user of the group, and / or a registered user can override the automatic generation of the device group 101B via a user interface input. However, when the device group 101B is automatically generated, the corresponding adaptation is made without first requiring a user interface input that explicitly indicates that the specific device group 101B is desired to be created (although an input at an earlier time can indicate approval to generally create the group). In some implementations, the device group 101B can be automatically generated in response to determining that the assistant devices 110A and 110B satisfy one or more proximity conditions with respect to each other. For example, the proximity conditions can include that the assistant devices 110A and 110B are assigned to the same structure (e.g., a particular house, a particular villa, a particular office) and / or the same room (e.g., a kitchen, a living room, a dining room) or other area within the same structure in the device topology. As another example, the proximity conditions can include that sensor signals from each of the assistant devices 110A and 110B indicate that they are proximally located to each other. For example, if both the assistant devices 110A and 110B consistently (e.g., more than 70% of the time or other threshold) detect the occurrence of wake words at the same time or a time close thereto (e.g., within 1 second), this can indicate that they are proximally located to each other. Also, for example, one of the assistant devices 110A and 110B can emit a signal (e.g., ultrasonic), and the other of the assistant devices 110A and 110B can attempt to detect the emitted signal.When the other of assistant devices 110A and 110B optionally detects the emitted signal by the strength of a threshold value, it may indicate that they are proximally located to each other. Additional and / or alternative techniques for determining temporal proximity and / or automatically generating a device group may be utilized.
[0047] Regardless of how Group 101B is generated, FIG. 1B1 shows an example of an adaptation that can be performed on assistant devices 110A and 110B in response to assistant devices 110A and 110B being included in Group 101B. In various implementations, one or both of assistant clients 120A and 120B can determine the adaptations to be performed and can cause those adaptations to occur. In other implementations, one or more engines of cloud-based assistant component 140 can, in addition to, or alternatively, determine the adaptations to be performed and can cause those adaptations to occur. As described herein, the adaptations to be performed can be determined based on considering the processing capabilities of both assistant clients 120A and 120B. For example, the adaptations can attempt to utilize the collective processing capabilities as much as possible while ensuring that the individual processing capabilities of each of the assistant devices are sufficient for the engines and / or models to be stored locally and utilized in the assistant. Further, the adaptations to be performed can also be determined based on usage data that reflects metrics related to the actual usage of the group of assistant devices and / or other non-grouped assistant devices in the ecosystem. For example, if the processing power enables either a larger but more accurate wake queue model, or a larger but more robust warm queue model, but not both, the usage data can be utilized to select one of the two options. For example, if the usage data reflects rare use (or no use at all) of warm words, and / or the detection of wake words barely exceeds the threshold frequently, and / or false negatives for wake words occur frequently, a larger but more accurate wake queue model can be selected. On the other hand, if the usage data reflects frequent use of warm words, and / or the detection of wake words consistently exceeds the threshold, and / or false negatives for wake words are rare, a larger but more robust warm queue model can be selected.The consideration of such usage data may determine whether the adaptation of FIG. 1B1, FIG. 1B2, or FIG. 1B3 is selected (FIG. 1B1, FIG. 1B2, or FIG. 1B3 each shows a different adaptation for the same group 101B).
[0048] In FIG. 1B1, the fulfillment engine 124A1 and the fulfillment model 134A1 and the TTS engine 125A1 and the TTS model 135A1 are purged from the assistant device 110A. Further, the assistant device 110A has a different ASR engine 122A2 and a different on-device ASR model 132A2, and further a different NLU engine 123A2 and a different on-device NLU model 133A2. The different engines can be downloaded from the local model repository 150, which is accessible in the assistant device 110A via interaction with the cloud-based assistant component 140. In FIG. 1B1, the wake queue engine 121B1, the ASR engine 122B1, and the authentication engine 126B1, and their corresponding models 131B1, 133B1, and 136B1 have already been purged from the assistant device 110B. Further, the assistant device 110B has a different NLU engine 123B2 and a different on-device NLU model 133B2, a different fulfillment engine 124B2 and a fulfillment model 133B2, and a different warmword engine 127B2 and a warmword model 137B2. The different engines can be downloaded from the local model repository 150, which is accessible in the assistant client 110B via interaction with the cloud-based assistant component 140.
[0049] The ASR engine 122A2 and the ASR model 132A2 of the assistant device 110A are more robust and / or accurate than the ASR engine 122A1 and the ASR model 132A1, but may occupy a large amount of disk space, utilize a large amount of memory, and / or require a large amount of processor resources. For example, the ASR model 132A1 can include only a single-pass model, and the ASR model 132A2 can include a two-pass model.
[0050] Similarly, the NLU engine 123A2 and the NLU model 133A2 are more robust and / or accurate than the NLU engine 123A1 and the NLU model 133A1, but may occupy a large amount of disk space, utilize a large amount of memory, and / or require a large amount of processor resources. For example, the NLU model 133A1 may already include only the intents and parameters for the first classification such as "lighting control", while the NLU model 133A2 may include the intents for not only "lighting control" but also "thermostat control", "smart lock control", and "reminders".
[0051] Therefore, the ASR engine 122A2, the ASR model 132A2, the NLU engine 123A2, and the NLU model 133A2 are improved with respect to the replaced counterparts. However, note that the processing capabilities of the assistant device 110A can prevent the ASR engine 122A2, the ASR model 132A2, the NLU engine 123A2, and the NLU model 133A2 from being stored and / or used without first purging the fulfillment engine 124A1 and the fulfillment model 134A1 and the TTS engine 125A1 and the TTS model 135A1. Simply purging such models from the assistant device 110A without complementary adaptation to and cooperative processing with the assistant device 110B would result in the assistant client 120A lacking the ability to process various assistant requests completely locally (i.e., without the need to utilize one or more cloud-based assistant components 140).
[0052] Accordingly, complementary adaptation is performed on the assistant device 110B, and after the adaptation, cooperative processing is performed between the assistant devices 110A and 110B. The NLU engine 123B2 and the NLU model 133B2 for the assistant device 110B are more robust and / or accurate than the NLU engine 123B1 and the NLU model 133B1, but may occupy a large amount of disk space, utilize a large amount of memory, and / or require a large amount of processor resources. For example, the NLU model 133B1 may already include intents and parameters for a first classification such as "lighting control". However, the NLU model 133B2 may cover more intents and parameters. Note that the intents and parameters covered by the NLU model 133B2 may be limited to intents that are not already covered by the NLU model 133A2 of the assistant client 120A. This can prevent functional duplication between the assistant clients 120A and 120B and extend the collective ability when the assistant clients 120A and 120B process assistant requests cooperatively.
[0053] Similarly, the fulfillment engine 124B2 and the fulfillment model 134B2 are more robust and / or accurate than the fulfillment engine 124B1 and the fulfillment model 124B1, but may occupy a large amount of disk space, utilize a large amount of memory, and / or require a large amount of processor resources. For example, the fulfillment model 124B1 may already only have the fulfillment ability for a single classification of the NLU model 133B1, but the fulfillment model 124B2 can include the fulfillment ability for all classifications of the NLU model 133B2 and even for the NLU model 133A2.
[0054] Accordingly, the fulfillment engine 124B2, the fulfillment model 134B2, the NLU engine 123B2, and the NLU model 133B2 are improved with respect to the replaced counterparts. However, the processing capacity of the assistant device 110B may prevent the fulfillment engine 124B2, the fulfillment model 134B2, the NLU engine 123B2, and the NLU model 133B2 from being stored and / or used without first purging the purged models and engines from the assistant device 110B. Simply purging such models from the assistant device 110B without complementary adaptation to and co-processing with the assistant device 110A will result in the assistant client 120B lacking the ability to fully process various assistant requests locally.
[0055] The warm queue engine 127B2 and the warm queue model 137B2 of the client device 110B do not occupy additional disk space, do not utilize much memory, and do not require significant processor resources compared to the warm queue engine 127B1 and the warm queue model 127B1. For example, the processing capabilities they require may be the same or even smaller. However, the warm queue engine 127B2 and the warm queue model 137B2 cover a warm queue in addition to what is covered by the warm queue engine 127B1 and the warm queue model 127B1 -- and in addition to what is covered by the warm queue engine 127A1 and the warm queue model 127A1 of the assistant client 120A.
[0056] In the configuration of FIG. 1B1, the assistant client 120A may be assigned processing roles of monitoring a wake queue, executing ASR, executing NLU for a first set of classifications, executing authentication, monitoring a first set of warm queues, and executing VAD. The assistant client 120B may be assigned processing roles of executing NLU for a second set of classifications, executing fulfillment, executing TTS, and monitoring a second set of warm queues. The processing roles may be communicated and stored in each of the assistant clients 120A, and the adjustment of the processing of various assistant requests may be executed by one or both of the assistant clients 120A and 120B.
[0057] As an example of cooperative processing of an assistant request using the adaptation of FIG. 1B1, an utterance of "OK Assistant, turn on the kitchen lights" is provided, and it is assumed that the assistant device 120A is the lead device that adjusts the processing. The wake queue engine 121A1 of the assistant client 120A can detect the appearance of the wake queue "OK Assistant". In response, the wake queue engine 121A1 can cause the ASR engine 122A2 to process the captured audio data following the wake queue. The wake queue engine 121A1 can optionally transmit a command to the assistant device 110B to transition from a low power state to a high power state locally to enable the assistant client 120B to immediately execute a specific processing of the assistant request. The audio data processed by the ASR engine 122A2 can be the audio data captured by the microphone of the assistant device 110A and / or the audio data captured by the microphone of the assistant device 110B. For example, the command transmitted to the assistant device 110B to transition to a higher power state can also cause it to locally capture audio data and optionally transmit such audio data to the assistant client 120A. In some implementations, the assistant client 120A can determine whether to use the received audio data or, instead, the audio data captured locally based on an analysis of the characteristics of each instance of the audio data. For example, an instance of audio data can be utilized over another instance based on having a lower signal-to-noise ratio and / or capturing an utterance at a higher volume.
[0058] The transcript generated by the ASR engine 122A2 is transmitted to the NLU engine 123A2 to perform NLU on the transcript for the first set of classifications, and can also be transmitted to the assistant client 120B to cause the NLU engine 123B2 to perform NLU on the transcript for the second set of classifications. The results of the NLU executed by the NLU engine 123B2 are transmitted to the assistant client 120A, and based on those results and the results from the NLU engine 123A2, it can be determined which results, if any, to utilize. For example, the assistant client 120A can utilize the result with the highest probability intent as long as its probability meets a certain threshold. For example, a result including the intent of "turn on" and a parameter specifying an identifier for "kitchen lights" can be utilized. It should be noted that if the probability does not meet the threshold, the NLU engine of the cloud-based assistant component 140 can optionally be utilized to perform NLU. The assistant client 120A can transmit the NLU result with the highest probability to the assistant client 120B. The fulfillment engine 124B2 of the assistant client 120B determines to transmit a command to transition "kitchen lights" to the "on" state using those NLU results, and such a command can be transmitted on the LAN 108. Optionally, the fulfillment engine 124B2 can utilize the TTS engine 125B1 to generate a synthetic voice to confirm the execution of "turning on the kitchen lights". In such a situation, the synthetic voice can be rendered by the assistant client 120B on the assistant device 110B and / or transmitted to the assistant device 110A for rendering by the assistant client 120A.
[0059] As another example of collaborative processing of assistant requests, assume that assistant client 120A is rendering an alarm of the local timer of assistant client 120A that has just finished. Further assume that the warm queue monitored by warm queue engine 127B2 contains "stop", and the warm queue monitored by warm queue engine 127A1 does not contain "stop". Finally, assume that when the alarm is being rendered, a spoken utterance of "stop" is provided and captured within the audio data detected via the microphone of assistant client 120B. Warm word engine 127B2 can process the audio data and determine the occurrence of the word "stop". Further, warm word engine 127B2 can determine that the occurrence of the word stop is directly mapped to a command to dismiss the timer or alarm that is sounding. That command is transmitted by assistant client 120B to assistant client 120A, thereby causing assistant client 120A to execute that command and dismiss the timer or alarm that is sounding. In some implementations, warm word engine 127B2 may only monitor for the occurrence of "stop" in some situations. In those implementations, assistant client 120A can transmit a command to cause warm word engine 127B2 to monitor for the occurrence of "stop" in response to or anticipating the rendering of the alarm. This command can cause monitoring to occur for a specific time period or, alternatively, until a stop monitoring command is sent.
[0060] Next, referring to FIG. 1B2, the same group 101B is illustrated. In FIG. 1B2, the same adaptation as in FIG. 1B1 is made, except that the ASR engine 121A1 and the ASR model 132A1 are not replaced by the ASR engine 122A2 and the ASR model 132A2. Rather, the ASR engine 121A1 and the ASR model 132A1 remain, and an additional ASR engine 122A3 and an additional ASR model 132A3 are provided.
[0061] The ASR engine 121A1 and the ASR model 132A1 may be for speech recognition of utterances in a first language (e.g., English), and the additional ASR engine 122A3 and the additional ASR model 132A3 may be for speech recognition of utterances in a second language (e.g., Spanish). The ASR engine 121A2 and the ASR model 132A2 in FIG. 1B1 may also be for the first language and may be more robust and / or accurate than the ASR engine 121A1 and the ASR model 132A1. However, the processing power of the assistant client 120A may prevent the ASR engine 121A2 and the ASR model 132A2 from being locally stored together with the ASR engine 122A3 and the additional ASR model 132A3. Nevertheless, this processing power enables the storage and use of both the ASR engine 121A1 and the ASR model 132A1 and the additional ASR engine 122A3 and the additional ASR model 132A3.
[0062] Instead of the ASR engine 121A2 and the ASR model 132A2 of FIG. 1B1, the decision to locally store both the ASR engine 121A1 and the ASR model 132A1 and the additional ASR engine 122A3 and the additional ASR model 132A3 may be based on usage statistics indicating that the vocal utterances provided by the assistant devices 110A and 110B (and / or the assistant devices 110C and 110D) in the example of FIG. 1B2 include both vocal utterances in the first language and vocal utterances in the second language. In the example of FIG. 1B1, the usage statistics show only the vocal utterances in the first language, and as a result, in FIG. 1B1, the more robust ASR engine 121A2 and ASR model 132A2 may be selected.
[0063] Referring next to FIG. 1B3, the same group 101B is illustrated again. In FIG. 1B3, the same adaptation as in the case of FIG. 1B1 is made, except that (1) instead of the ASR engine 121A1 and the ASR model 132A1 being replaced by the ASR engine 122A2 and the ASR model 132A2, they are replaced by the ASR engine 122A4 and the ASR model 132A4, (2) there is no warm queue engine or warm queue model on the assistant device 110B, and (3) the ASR engine 122B4 and the ASR model 132B4 are locally stored and used on the assistant device 110B.
[0064] The ASR engine 122A4 and the ASR model 132A4 are used to execute the first part of speech recognition, and the ASR engine 122B4 and the ASR model 132B4 can be used to execute the second part of speech recognition. For example, the ASR engine 122A4 uses the ASR model 132A4 to generate an output, and the output is transmitted to the assistant client 120B. The ASR engine 122B4 can process the output when generating the recognition of the voice. As one specific example, the output is a graph representing candidate recognitions, and the ASR engine 122B4 can perform beam search on the graph when generating the recognition of the voice. As another specific example, the ASR model 132A4 can be the initial / downstream part (i.e., the first neural network layer) of the end-to-end speech recognition model, and the ASR model 132B4 can be the later / upstream part (i.e., the second neural network layer) of the end-to-end speech recognition model. In such an example, the end-to-end model is split between the two assistant devices 110A and 110B, and the output can be in the state of the last layer (e.g., embedding) of the initial part after processing. As yet another example, the ASR model 132A4 can be an acoustic model, and the ASR model 132B4 can be a language model. In such an example, the output can indicate a sequence of phonemes or a sequence of probability distributions for phonemes, and the ASR engine 122B4 can use the language model to select the transcript / recognition corresponding to that sequence.
[0065] The robustness and / or accuracy of the ASR engines 122A4, ASR models 132A4, ASR engines 122B4, and ASR models 132B4 operating in coordination can exceed that of the ASR engine 122A2 and ASR model 132A2 of FIG. 1B1. Further, the processing capabilities of assistant clients 120A and 120B can prevent the ASR models 132A4 and 132B4 from being individually stored and utilized on any of those devices. However, the processing capabilities can enable the splitting of the models and the splitting of the processing roles between the ASR engines 122A4 and 122B4, as described herein. Note that on assistant device 110B, purging of the warm queue engine and warm queue models can enable the storage and utilization of the ASR engine 122B4 and ASR model 132B4. In other words, the processing capabilities of this assistant device 110B will not enable the storage and / or utilization of the warm queue engine and warm queue models, along with the other engines and models illustrated in FIG. 1B3.
[0066] Instead of the ASR engine 121A2 and ASR model 132A2 of FIG. 1B1, the decision to locally store the ASR engines 122A4, ASR models 132A4, ASR engines 122B4, and ASR models 132B4 may be based on usage statistics indicating that speech recognition on assistant devices 110A and 110B (and / or assistant devices 110C and 110D) is often of low confidence and / or often inaccurate in the example of FIG. 1B3. For example, the usage statistics may indicate that the confidence metric for recognition is below average (e.g., an average based on a population of users) and / or that the recognition is often corrected by the user (e.g., through editing of the displayed transcript).
[0067] Next, referring to FIG. 1C, a device group 101C is created, and assistant devices 110A, 110B, and 110C are included in the device group 101C. In some implementations, the device group 101C can be generated in response to a user interface input that explicitly indicates a desire to group the assistant devices 110A, 110B, and 110C. For example, the user interface input can indicate a desire to create the device group 101C from scratch, or alternatively, to add the assistant device 110C to the device group 101B (FIGS. 1B1, 1B2, and 1B3), thereby creating a modified group 101C. In some implementations, the device group 101C can be generated automatically instead. For example, the device group 101B (FIGS. 1B1, 1B2, and 1B3) may have been previously generated based on determining that the assistant devices 110A and 110B are in close proximity, and after the creation of the device group 101B, the assistant device 110C can be moved by the user to be proximal to the devices 110A and 110B. As a result, the assistant device 110C is automatically added to the device group 101B, thereby creating a modified group 101C.
[0068] Regardless of how the group 101C is generated, FIG. 1C shows an example of an adaptation that can be performed on the assistant devices 110A, 110B, and 110C in response to the assistant devices 110A, 110B, and 110C being included in the group 101C.
[0069] In FIG. 1C, the assistant device 110B has the same adaptation as in the case of FIG. 1B3. Further, the assistant device 110A has the same adaptation as in the case of FIG. 1B3, except that (1) the authentication engine 126A1 and the VAD engine 128A1, and their corresponding models 13A1 and 138A1 have been purged, (2) the wake queue engine 121A1 and the wake queue model 131A1 have been replaced with the wake queue engine 121A2 and the wake queue model 131A2, and (3) the warm queue engine or the warm queue model does not exist on the assistant device 110B. The models and engines stored on the assistant device 110C are not adapted. However, the assistant client 120C can be adapted to enable cooperative processing of assistant requests with the assistant clients 120A and 120B.
[0070] In FIG. 1C, the authentication engine 126A1 and the VAD engine 128A1 have been purged from the assistant device 110A because their counterparts already exist on the assistant device 110C. In some implementations, the authentication engine 126A1 and / or the VAD engine 128A1 may be purged only after some or all of their data from those components has been merged with their counterparts that already exist on the assistant device 110C. As an example, the authentication engine 126A1 can store voice embeddings for a first user and a second user, while the authentication engine 126C1 can store only voice embeddings for the first user. Before purging the authentication engine 126A1, the voice embeddings for the second user are locally transmitted to the authentication engine 126C1, thereby ensuring that such voice embeddings are utilized by the authentication engine 126C1 and that the pre-adaptation capabilities are maintained after adaptation. As another example, the authentication engine 126A1 can capture utterances of a second user and store instances of voice data utilized to generate voice embeddings for the second user, while the authentication engine 126C1 may lack voice embeddings for the second user. Before purging the authentication engine 126A1, the instances of voice data are locally transmitted from the authentication engine 126A1 to the authentication engine 126C1, thereby ensuring that the instances of voice data are utilized by the authentication engine 126C1 to generate voice embeddings for the second user using the on-device authentication model 136C1 and that the pre-adaptation capabilities are maintained after adaptation. Further, the wake queue engine 121A1 and the wake queue model 131A1 have been replaced by the wake queue engine 121A2 and the wake queue model 131A2 with a smaller storage size. For example, the wake queue engine 121A1 and the wake queue model 131A1 enabled detection of both a vocal wake queue and an image-based wake queue, while the wake queue engine 121A2 and the wake queue model 131A2 enable detection of only an image-based wake queue.Optionally, personalizations, training instances, and / or other settings from the image-based wake queue portions of the wake queue engine 121A1 and wake queue model 131A1 may be merged with or otherwise shared with the wake queue engine 121A2 and wake queue model 131A2 prior to purging the wake queue engine 121A1 and wake queue model 131A1. The wake queue engine 121C1 and wake queue model 131C1 are capable of detecting only the vocal wake queue. Thus, the wake queue engine 121A2 and wake queue model 131A2, and the wake queue engine 121C1 and wake queue model 131C1, collectively enable detection of both vocal and image-based wake queues. Optionally, personalizations and / or other settings from the vocal queue portions of the wake queue engine 121A1 and wake queue model 131A1 may be transmitted to the client device 110C for merging with or otherwise sharing with the wake queue engine 121C1 and wake queue model 131C1.
[0071] Further, extra storage space is obtained by replacing the wake queue engine 121A1 and wake queue model 131A1 with the wake queue engine 121A2 and wake queue model 131A2 of a smaller storage size. The extra storage space resulting from purging the extra storage space, as well as the authentication engine 126A1 and VAD engine 128A1, and their corresponding models 13A1 and 138A1, provides space for the warm queue engine 127A2 and warm queue model 137A2 (which is collectively larger than the replaced warm queue engine 127A1 and warm queue model). The warm queue engine 127A2 and warm queue model 137A2 may be utilized to monitor a warm queue different from that monitored using the warm queue engine 127C1 and warm queue model 137C1.
[0072] As an example of cooperative processing of an assistant request using the adaptation of FIG. 1C, an utterance of "OK Assistant, turn on the kitchen lights" is provided, and it is assumed that the assistant device 120A is the lead device that adjusts the processing. The wake queue engine 121C1 of the assistant client 110C can detect the appearance of the wake queue "OK Assistant". In response to this, the wake queue engine 121C1 can transmit a command to the assistant devices 110A and 110B to cause them to cooperatively process the captured audio data following the wake queue by the ASR engines 122A4 and 122B4. The audio data to be processed may be the audio data captured by the microphone of the assistant device 110C, and / or the audio data captured by the microphone of the assistant device 110B and / or by the assistant device 110C.
[0073] The transcript generated by the ASR engine 122B4 is transmitted to the NLU engine 123B2 to perform NLU on the transcript for the second set of classifications, and can also be transmitted to the assistant client 120A to cause the NLU engine 123A2 to perform NLU on the transcript for the first set of classifications. The results of the NLU performed by the NLU engine 123B2 are transmitted to the assistant client 120A, and based on those results and the results from the NLU engine 123A2, it can be determined which results, if any, to utilize. The assistant client 120A can transmit the NLU result with the highest probability to the assistant client 120B. The fulfillment engine 124B2 of the assistant client 120B determines to transmit a command to transition "kitchen lights" to the "on" state using those NLU results to the "kitchen lights", and such a command can be transmitted on the LAN 108. Optionally, the fulfillment engine 124B2 can utilize the TTS engine 125B1 to generate a synthetic voice to confirm the execution of "turning on the kitchen lights". In such a situation, the synthetic voice can be rendered by the assistant client 120B on the assistant device 110B, transmitted to the assistant device 110A for rendering by the assistant client 120A, and / or transmitted to the assistant device 110C for rendering by the assistant client 120C.
[0074] Next, referring to FIG. 1D, a device group 101D is created, and the assistant devices 110C and 110D are included in the device group 101D. In some implementations, the device group 101D can be generated in response to a user interface input that explicitly indicates the desire to group the assistant devices 110C and 110D. In some implementations, the device group 101D can alternatively be generated automatically.
[0075] Regardless of how Group 101D is generated, FIG. 1D shows an example of an adaptation that can be performed on assistant devices 110C and 110D in response to the assistant devices 110C and 110D being included in Group 101D.
[0076] In FIG. 1D, the wake queue engine 121C1 and the wake queue model 131C1 of assistant device 110C are replaced by the wake queue engine 121C2 and the wake queue model 131C2. Further, the authentication engine 126D1 and the authentication model 136D1, as well as the VAD engine 128D1 and the VAD model 138D2, are purged from the assistant device 110D. Even further, the wake queue engine 121D1 and the wake queue model 131D1 of the assistant device 110D are replaced by the wake queue engine 121D2 and the wake queue model 131D2, and the warm queue engine 127D1 and the warm queue model 137D1 are replaced by the warm queue engine 127D2 and the warm queue model 137D2.
[0077] In the case of the previous wake queue engine 121C1 and wake queue model 131C1, it could only be used to detect a first set of one or more wake words such as "Hey Assistant" and "OK Assistant". On the other hand, the wake queue engine 121C2 and wake queue model 131C2 can only detect an alternative second set of one or more wake words such as "Hey Computer" and "OK Computer". Also, in the case of the previous wake queue engine 121D1 and wake queue model 131D1, it could only have been used to detect a first set of one or more wake words, and the wake queue engine 121D2 and wake queue model 131D2 can only be used to detect a first set of one or more wake words. However, the wake queue engine 121D2 and wake queue model 131D2 are larger than the replaced counterparts and have higher robustness (e.g., more robust to background noise) and / or accuracy. Purging the engine and model from the assistant device 110D may enable the use of the larger-sized wake queue engine 121D2 and wake queue model 131D2. Further, collectively, the wake queue engine 121C2 and wake queue model 131C2 and the wake queue engine 121D2 and wake queue model 131D2 enable the detection of two sets of wake words, but each of the assistant clients 120C and 120D was only capable of detecting the first set before adaptation.
[0078] The warm queue engine 127D2 and the warm queue model 137D2 of the assistant device 110D may require more computing power than the replaced wake queue engine 127D1 and wake queue model 137D1. However, these capabilities are available through purging the engines and models from the assistant device 110D. Further, the warm queues monitored by the warm queue engine 127D2 and the warm queue model 137D1 can be added to those monitored by the warm queue engine 127C1 and the warm queue model 137D1. Prior to adaptation, the wake queues monitored by the wake queue engine 127D1 and the wake queue model 137D1 were the same as those monitored by the warm queue engine 127C1 and the warm queue model 137D1. Thus, through cooperative processing, the assistant clients 120C and 120D can monitor more wake queues.
[0079] Note in the example of FIG. 1D that there are many assistant requests that cannot be fully processed on-device cooperatively by the assistant clients 120C and 120D. For example, the assistant clients 120C and 120D lack an ASR engine, lack an NLU engine, and lack a fulfillment engine. This may be due to the processing capabilities of the assistant devices 110C and 110D not being able to support any such engines or models. Thus, for voice utterances that are not warm queues supported by the assistant clients 120C and 120D, the cloud-based assistant component 140 needs to be used at all times when fully processing many assistant requests. However, the adaptation and cooperative processing based on the adaptation of FIG. 1D can still be more robust and / or accurate compared to any processing performed individually on the device prior to adaptation. For example, the adaptation enables the detection of additional wake queues and additional warm queues.
[0080] As an example of collaborative processing that can be performed, assume the spoken utterance "OK Computer, play some music". In such an example, the wake queue engine 121D2 can detect the wake queue "OK Computer". In response, the wake queue engine 121D2 can cause the voice data corresponding to the wake queue to be transmitted to the assistant client 120C. The authentication engine 126C1 of the assistant client 120C can use the voice data to determine whether the utterance of the wake queue can be authenticated for the registered user. The wake queue engine 121D2 can further stream the voice data following the spoken utterance to the cloud-based assistant component 140 for further processing. The voice data can be captured at the assistant device 110D or at the assistant device 110C (for example, the assistant client 120 can transmit a command to cause the assistant client 120C to capture the voice data in response to the wake queue engine 121D2 detecting the wake queue). Further, authentication data based on the output of the authentication engine 126C1 can also be transmitted along with the voice data. For example, when the authentication engine 126C1 authenticates the utterance of the wake queue for the registered user, the authentication data can include an identifier of the registered user. As another example, when the authentication engine 126C1 does not authenticate the utterance of the wake queue for any registered user, the authentication data can include an identifier reflecting the utterance provided by the guest user.
[0081] Various specific examples have been described so far with reference to FIGS. 1B1, 1B2, 1B3, 1C, and 1D. However, it should be noted that various additional or alternative groups can be generated and / or various additional or alternative adaptations can be performed in response to the generation of the groups.
[0082] FIG. 2 is a flowchart illustrating an exemplary method 200 for adapting an on-device model and / or processing role of assistant devices within a group. For convenience, the operations of the flowchart are described with reference to a system that performs the operations. The system can include various components of various computer systems, such as one or more of the assistant clients 120A-D of FIG. 1 and / or components of the cloud-based assistant component 140 of FIG. 1. Further, although the operations of method 200 are shown in a particular order, this is not meant to be limiting. One or more operations may be reordered, omitted, or added.
[0083] In block 252, the system generates a group of assistant devices. For example, the system can generate a group of assistant devices in response to a user interface input that explicitly indicates a desire to generate a group. As another example, the system can automatically generate a group in response to determining that one or more conditions are met. As yet another example, the system can automatically determine that a group should be generated in response to determining that a condition is met, provide a user interface output that suggests generating the group, and then generate the group in response to receiving an affirmative user interface in response to the user interface output.
[0084] At block 254, the system obtains the processing capabilities for each of the group's assistant devices. For example, the system may be one of the group's assistant devices. In such an example, the assistant device can obtain its own processing capabilities, and the other assistant devices in the group can communicate that processing capability to the assistant device. As another example, the processing capabilities of the assistant devices are stored in the device topology, and the system can retrieve them from the device topology. As yet another example, the system may be a cloud-based component, and each of the group's assistant devices can communicate its processing capabilities to the system.
[0085] The processing capabilities of the assistant devices can include corresponding processor values based on the capabilities of one or more on-device processors, corresponding memory values based on the size of the on-device memory, and / or corresponding disk space values based on the available disk space. For example, the processor value can include details regarding one or more operating frequencies of the processor, details regarding the size of the processor's cache, whether each of the processors is a GPU, CPU, or DSP, and / or other details. As another example, the processor value can, in addition to or alternatively, include a higher-level classification of the processor's capabilities such as high, medium, or low, or GPU+CPU+DSP, high-output CPU+DSP, medium-output CPU+DSP, or low-output CPU+DSP. As another example, the memory value can include details of the memory such as a specific size of the memory, or a higher-level classification of the memory such as high, medium, or low. As yet another example, the disk space value can include details regarding the available disk space such as a specific size of the disk space, or a higher-level classification of the available disk space such as high, medium, or low.
[0086] In block 256, the system utilizes the processing power of block 254 when determining a collective set of on-device models for a group. For example, the system may determine a set of on-device models that attempts to maximize the use of collective processing power, ensuring that each on-device model in the set can be locally stored and utilized on a device that can store and utilize the on-device model. The system may also attempt to ensure, if possible, that the selected set includes a complete (or more complete than other candidate sets) pipeline of on-device models. For example, the system may select a set that includes an ASR model over a set that includes a highly robust NLU model but no ASR model.
[0087] In some implementations, block 256 includes sub-block 256A, where the system uses usage data when selecting a collective set of on-device models for a group. Past usage data may be data related to past assistant interactions in one or more assistant devices of the group and / or in one or more additional assistant devices of the ecosystem. In some implementations, in sub-block 256A, the system considers usage data along with the above considerations when selecting the on-device models to include in the set. For example, if the processing power allows including a high-accuracy ASR model (instead of a low-accuracy ASR model) or a highly robust NLU model (instead of a less robust NLU model) in the set, but does not allow including both, the usage data can be used to determine which to select. For example, if the usage data reflects that past assistant interactions were predominantly (or exclusively) directed at intents covered by a less robust NLU model that can be included in the set with a high-accuracy ASR model, the high-accuracy ASR model can be selected to be included in the set. On the other hand, if the usage data reflects that past assistant interactions cover many intents covered by a highly robust NLU model but not by a less robust NLU model, the highly robust NLU model can be selected to be included in the set. In some implementations, the candidate set is first determined based on processing power without considering usage data, and then, if there are multiple valid candidate sets, the usage data can be used to select one over the other.
[0088] In block 258, the system causes each of the assistant devices to locally store a corresponding subset of the collective set of on-device models. For example, the system can communicate to each of the group of assistant devices the corresponding instructions as to which on-device models should be downloaded. Based on the received instructions, each of the assistant devices can download the corresponding models from the remote database. As another example, the system can retrieve the on-device models and push the corresponding on-device models to each of the group of assistant devices. As yet another example, for any on-device model that is to be stored on a corresponding one of the group of assistant devices prior to adaptation and on a corresponding other of the group of assistant devices upon adaptation, such model can be communicated directly between the respective devices. For example, assume that a first assistant device stores an ASR model prior to adaptation and that same ASR model is stored on a second assistant device upon adaptation and purged from the first assistant device. In such a case, the system instructs the first assistant device to transmit the ASR model to the second assistant device to the local storage of the second assistant device (and / or instructs the second assistant device to download from the first assistant device), and the first assistant device can then purge the ASR model. In addition to avoiding WAN traffic, transmitting the pre-adaptation models locally can maintain any personalization of those on-device models that was previously done at the time of transmission. The personalized models can have higher accuracy for the users of the ecosystem compared to the non-personalized counterparts in remote storage.As yet another example, for any assistant device that includes, prior to adaptation, stored training instances for personalizing an on-device model on those assistant devices, such training instances can be communicated, after adaptation, to an assistant device that will have a corresponding model downloaded from a remote database. An assistant device that has an on-device model after adaptation can then use the training instances to personalize the corresponding model downloaded from the remote database. The corresponding model downloaded from the remote database may be different (e.g., smaller or larger) than the counterpart that the training instances were used on prior to adaptation, but the training instances can be used as-is when personalizing different downloaded on-device models.
[0089] In block 260, the system assigns corresponding roles to each of the assistant devices. In some implementations, assigning the corresponding roles may include causing each of the assistant devices to download and / or implement an engine corresponding to the on-device model stored locally on the assistant device. Each engine can utilize the corresponding on-device model when performing corresponding processing roles such as performing all or some of the ASR, performing wake word recognition for at least some wake words, performing warm word recognition for some warm words, and / or performing authentication. In some implementations, one or more of the processing roles are executed only when there is a command from the lead device of a group of assistant devices. For example, the NLU processing role performed by a given device using an on-device NLU model may be executed only in response to the lead device transmitting corresponding text for NLU processing and / or a specific command to cause the given device to perform the NLU processing. As another example, the warm word monitoring processing role performed by a given device using an on-device warm queue engine and an on-device warm queue model may be executed only in response to the lead device transmitting a command to cause the given device to perform warm word processing. For example, the lead device can cause a given device to monitor for the vocal occurrence of the "stop" warm word in response to an alarm sounding on the lead device or another device in the group. In some implementations, one or more of the processing roles can be executed at least selectively independent of any command from the lead assistant device. For example, the wake queue monitoring role performed by a given device using an on-device wake queue engine and an on-device wake queue model can be continuously executed unless explicitly disabled by the user.As another example, a warm queue monitoring role performed by a given device using an on-device warm queue engine and an on-device warm queue model may be executed continuously or based on conditions where monitoring is detected locally at the given device.
[0090] At block 262, the system causes subsequent utterances detected at one or more of the group of devices to be cooperatively and locally processed at the group's assistant device according to its role. Various non-limiting examples of such cooperative processing are described herein. For example, the examples are described with reference to FIGS. 1B1, 1B2, 1B3, 1C, and 1D.
[0091] At block 264, the system determines whether there has been any change to the group, such as adding a device to the group, removing a device from the group, or dissolving the group. If not, the system continues to execute block 262. If so, the system proceeds to block 266.
[0092] At block 266, the system determines whether the change to the group has left one or more of the assistant devices that were within the group now alone (i.e., no longer assigned to the group). If so, the system proceeds to block 268, where each of the standalone devices is caused to locally store the pre-grouping on-device model and assume its pre-grouping on-device processing role. In other words, a device may be returned to the state it was in prior to adaptation being executed in response to being included in the group if it is no longer within the group. In these and other ways, after returning to that state, the standalone device can functionally process various assistant requests operating in its standalone capacity. Prior to returning to that state, the standalone device may not have been able to functionally process any assistant requests, or at least a lesser amount of assistant requests than prior to returning to that state.
[0093] In block 270, the system determines whether two or more devices remain in the modified group. If so, the system returns to block 254 and performs another iteration of blocks 254, 256, 258, 260, and 262 based on the modified group. For example, if the modified group includes additional assistant devices without losing any of the previous assistant devices of the group, adaptation can be made considering the additional processing capabilities of the additional assistant devices. If the determination at block 270 is no, the group is dissolved and the system proceeds to block 272 where method 200 ends (until another group is generated).
[0094] FIG. 3 is a flowchart illustrating an exemplary method 300 that may be performed by each of a plurality of assistant devices within a group when adapting the on-device models and / or processing roles of the assistant devices within the group. Although the operations of method 300 are shown in a particular order, this is not meant to be limiting. One or more operations may be reordered, omitted, or added.
[0095] The operations of method 300 are a particular example of method 200 that may be executed by each of the assistant devices within a group. Accordingly, the operations are described with reference to the assistant device(s) performing the operations, such as one or more of assistant clients 120A - D of FIG. 1. Each of the assistant devices within the group can execute method 300 in response to receiving an input indicating that it is included in the group.
[0096] In block 352, the assistant device receives a grouping indication indicating that it was included in a group. In block 352, the assistant device also receives identifiers of other assistant devices within the group. Each identifier can be, for example, a MAC address, an IP address, a label assigned to the device (e.g., the use assigned in the device topology), a serial number, or other identifier.
[0097] In optional block 354, the assistant device transmits data to other assistant devices in the group. The data is transmitted to other devices using the identifiers received in block 352. Put another way, the identifier can be a network address or can be used to find the network address that is the destination of the data. The transmitted data can include one or more of the processing values described herein, another device identifier, and / or other data.
[0098] In optional block 356, the assistant device receives data transmitted by other devices in block 354.
[0099] At block 358, the assistant device determines whether it is a leader device based on the optionally received data at block 356 or the identifier received at block 352. For example, the device can select itself as the leader if its own identifier is the lowest value (or alternatively the highest value) compared to the other identifiers received at block 352. As another example, the device can select itself as the leader if its processed value exceeds the processed value received within the data at optionally block 356 among all other processed values. Other data may be transmitted at block 354 and received at block 356, and such other data can similarly enable an objective determination in the assistant device as to whether it should be the leader. More generally, at block 358, the assistant device can utilize one or more objective criteria when determining whether it should be the leader device.
[0100] At block 360, the assistant device determines whether it was determined at block 358 to be the leader device. An assistant device that was not determined to be the leader device then branches down from block 360 to the "no" branch. An assistant device that was determined to be the leader device branches down from block 360 to the "yes" branch.
[0101] In the "yes" branch, the assistant device, at block 360, utilizes the processing capabilities received from other assistant devices within the group, as well as its own processing capabilities, when determining the collective set of on-device models for the group. The processing capabilities can be transmitted by other assistant devices to the lead device at optional block 354 or block 370 (described later) when optional block 354 is not executed or the data of block 354 does not include the processing capabilities. In some implementations, block 362 can share one or more aspects common to block 256 of method 200 in FIG. 2. For example, in some implementations, block 362 can also include considering past usage data when determining the collective set of on-device models.
[0102] At block 364, the assistant device transmits to each of the other assistant devices in the group the respective instructions of the collective set of on-device models that the other assistant devices should download. Optionally, at block 364, the assistant device can also transmit to each of the other assistant devices in the group the respective instructions of the processing roles to be executed by the assistant device using the on-device models.
[0103] At block 366, the assistant device downloads and stores the set of on-device models assigned to the assistant device. In some implementations, blocks 364 and 366 can share one or more aspects common to block 258 of method 200 in FIG. 2.
[0104] At block 368, the assistant device coordinates the cooperative processing of the assistant request, including using its own on-device model when executing a part of the cooperative processing. In some implementations, block 368 can share one or more aspects common to block 262 of method 200 in FIG. 2.
[0105] Next, turning to the "no" branch, at optional block 370, the assistant device communicates its processing capabilities to the lead device. Block 370 may be omitted, for example, when block 354 is executed and the processing capabilities are included in the data transmitted at block 354.
[0106] At block 372, the assistant device receives from the lead device an indication of the on-device model to be downloaded and, optionally, an indication of the processing role.
[0107] At block 374, the assistant device downloads and stores the on-device model reflected within the indication of the on-device model received at block 372. In some implementations, blocks 372 and 374 can share one or more aspects common to block 258 of method 200 of FIG. 2.
[0108] At block 376, the assistant device utilizes its on-device model when performing a portion of the collaborative processing of the assistant request. In some implementations, block 376 can share one or more aspects common to block 262 of method 200 of FIG. 2.
[0109] FIG. 4 is a block diagram of an exemplary computing device 410 that may optionally be utilized to implement one or more aspects of the techniques described herein. In some implementations, one or more of the assistant device, and / or other components, may include one or more components of the exemplary computing device 410.
[0110] Computing device 410 typically includes at least one processor 414 that communicates with a number of peripheral devices via a bus subsystem 412. These peripheral devices may include, for example, a storage subsystem 425, including a memory subsystem 425 and a file storage subsystem 426, a user interface output device 420, a user interface input device 422, and a network interface subsystem 416. The input and output devices enable interaction between the user and the computing device 410. The network interface subsystem 416 provides an interface to an external network and is coupled to a corresponding interface device within other computing devices.
[0111] The user interface input device 422 may include, for example, a keyboard, a mouse, a trackball, a pointing device such as a touchpad or a graphics tablet, a scanner, a touch screen incorporated into a display, a voice input device such as a voice recognition system, a microphone, and / or other types of input devices. In general, the use of the term "input device" is intended to include all possible types of devices and methods for inputting information into the computing device 410 or onto a communication network.
[0112] The user interface output device 420 may include a non-visual display such as a display subsystem, a printer, a fax machine, or an audio output device. The display subsystem may include a cathode ray tube ("CRT"), a flat panel device such as a liquid crystal display ("LCD"), a projection device, or any other mechanism for generating a visible image. The display subsystem may also include a non-visual display via an audio output device or the like. In general, the use of the term "output device" is intended to include all possible types of devices and methods for outputting information from the computing device 410 to the user or another machine or computing device.
[0113] The storage subsystem 425 stores programming and data structures that implement some or all of the functions of some of the modules described herein. For example, the storage subsystem 425 may include logic circuitry for executing one or more selected aspects of the methods described herein and / or implementing the various components depicted herein.
[0114] These software modules are generally executed by the processor 414, either alone or in combination with other processors. The memory 425 used in the storage subsystem 425 can include a number of memories, such as a main random access memory ("RAM") 430 for storing instructions and data during program execution, and a read-only memory ("ROM") 432 in which fixed instructions are stored. The file storage subsystem 426 can provide a persistent storage area for program and data files, and may include, for example, a hard disk drive, a floppy disk drive with an associated removable medium, a CD-ROM drive, an optical drive, or a removable media cartridge. The modules implementing the functions of some implementations can be stored within the storage subsystem 425 by the file storage subsystem 426, or within other machines accessible by the processor 414.
[0115] The bus system 412 provides a mechanism for communicating between the various components and subsystems of the computing device 410 as intended. Although the bus system 412 is schematically illustrated as a single bus, alternative implementations of the bus system may use multiple buses.
[0116] The computing device 410 can be various types of devices, including a workstation, a server, a computing cluster, a blade server, a server farm, or any other data processing system or computing device. Since computers and networks are constantly changing in nature, the description of the computing device 410 shown in FIG. 4 is intended only as a specific example for the purpose of illustrating some implementations. Many other configurations of the computing device 410 are possible that have more or fewer components than the computing device shown in FIG. 4.
[0117] In situations where the systems described in this specification may collect or use personal information about a user (or what is often referred to as a "participant" in this specification), the user may be provided with the opportunity to control whether a program or function collects user information (e.g., information about the user's social network, social behavior or activities, professional occupation, user preferences, or the user's current geographical location), or to control whether and / or how the user receives content from a content server that is deemed to be of higher relevance to the user. Also, specific data may be processed in one or more ways before it is stored or used, and thus personally identifiable information is removed. For example, the user's identity may be processed so that personally identifiable information cannot be determined for the user, or the user's geographical location may be generalized when geographical location information (such as city name, postal code, country level, etc.) is obtained, such that the user's specific geographical location need not be determined. Thus, the user may control how information about the user is collected and / or used.
[0118] In some implementations, a method is provided that includes generating an assistant device group of heterogeneous assistant devices. The heterogeneous assistant devices include at least a first assistant device and a second assistant device. At the time of generating the group, the first assistant device includes a first set of on-device models that are locally stored and used when locally processing assistant requests directed to the first assistant device. Further, at the time of generating the group, the second assistant device includes a second set of on-device models that are locally stored and used when locally processing assistant requests directed to the second assistant device. The method further includes determining a collective set of on-device models that are locally stored for use when cooperatively and locally processing assistant requests directed to any of the heterogeneous assistant devices in the assistant device group, based on the respective processing capabilities of the heterogeneous assistant devices in the assistant device group. The method further includes, in response to the generation of the assistant device group, locally storing in each of the heterogeneous assistant devices a corresponding subset of the collective set of on-device models that are locally stored, and assigning one or more corresponding processing roles to each of the heterogeneous assistant devices in the assistant device group. Each of the processing roles utilizes one or more corresponding ones of the on-device models that are locally stored. Further, locally storing in each of the heterogeneous assistant devices a corresponding subset includes purging one or more first on-device models of the first set from the first assistant device to provide storage space for the corresponding subset that is locally stored on the first assistant device, and purging one or more second on-device models of the second set from the second assistant device to provide storage space for the corresponding subset that is locally stored on the second assistant device.This method further includes, following the assignment of corresponding processing roles to each of the heterogeneous assistant devices of the assistant device group, detecting an uttered speech via a microphone of at least one device of the heterogeneous assistant devices of the assistant device group, and causing the uttered speech to be processed locally and cooperatively by the heterogeneous assistant devices of the assistant device group using its corresponding processing role in response to the uttered speech being detected via the microphone of the assistant device group.
[0119] These and other implementations of the techniques disclosed herein can optionally include one or more of the following features.
[0120] In some implementations, purging one or more first on-device models of a first set from a first assistant device includes purging a wake word detection model of a first set of a first device, which is used when detecting a first wake word, from the first assistant device. In those implementations, a corresponding subset locally stored on a second assistant device includes a wake word detection model of a second device that is used when detecting the first wake word, and assigning a corresponding processing role includes assigning a first wake word detection role to the second assistant device to use the wake word detection model of the second device when monitoring for the occurrence of the first wake word. In some of those implementations, an utterance includes a first wake word followed by an assistant command, and in the first wake word detection role, the second assistant device detects the occurrence of the first wake word and triggers the execution of an additional role of the corresponding processing role in response to the detection of the occurrence of the first wake word. In some versions of those implementations, the additional role of the corresponding processing role is executed by the first assistant device, and the second assistant device triggers the execution of the additional role of the corresponding processing role by transmitting an indication of the detection of the first wake word to the first assistant device.
[0121] In some implementations, the corresponding subset stored locally on the first assistant device includes the first wake word detection model of the first device used when detecting one or more first wake words, and excludes the wake word detection model used when detecting one or more second wake words. In some of those implementations, the corresponding subset stored locally on the second assistant device includes the second wake word detection model of the second device used when detecting one or more second wake words, and excludes the wake word detection model used when detecting one or more first wake words. In some versions of those implementations, assigning the corresponding processing roles includes assigning a first wake word detection role to the first assistant device to use the wake word detection model of the first device when monitoring for the occurrence of one or more first wake words, and assigning a second wake word detection role to the second assistant device to use the wake word detection model of the second device when monitoring for the occurrence of one or more second wake words.
[0122] In some implementations, the corresponding subset stored locally on the first assistant device includes a first language speech recognition model that is used when performing speech recognition in the first language and excludes any speech recognition models that are used when performing speech recognition in the second language. In some of those implementations, the corresponding subset stored locally on the second assistant device includes a second language speech recognition model that is used when performing speech recognition in the second language and excludes any speech recognition models that are used when performing speech recognition in the second language. In some versions of those implementations, assigning the corresponding processing roles includes assigning to the first assistant device a first language speech recognition role that uses the first language speech recognition model when performing speech recognition in the first language and assigning to the second assistant device a second language speech recognition role that uses the second language speech recognition model when performing speech recognition in the second language.
[0123] In some implementations, the corresponding subset stored locally on the first assistant device includes a first portion of a speech recognition model that is utilized when performing a first portion of speech recognition, and excludes a second portion of the speech recognition model. In some of those implementations, the corresponding subset stored locally on the second assistant device includes a second portion of the speech recognition model that is utilized when performing a second portion of speech recognition, and excludes a first portion of the speech recognition model. In some versions of those implementations, assigning the corresponding processing roles includes assigning a first portion of a language speech recognition role that utilizes the first portion of the speech recognition model to the first assistant device when generating a corresponding embedding for a corresponding speech, transmitting the corresponding embedding to the second assistant device, and assigning a second language speech recognition role that utilizes the corresponding embedding from the first assistant device and a second language speech recognition model when generating a corresponding recognition of the corresponding speech to the second assistant device.
[0124] In some implementations, the corresponding subset stored locally on the first assistant device includes a speech recognition model that is utilized when performing a first portion of speech recognition. In some of those implementations, assigning the corresponding processing roles includes assigning a first portion of a language speech recognition role that utilizes the speech recognition model when generating an output to the first assistant device, transmitting the corresponding output to the second assistant device, and assigning a second language speech recognition role that performs beam search with the corresponding output from the first assistant device when generating a corresponding recognition of the corresponding speech to the second assistant device.
[0125] In some implementations, the corresponding subset stored locally on the first assistant device includes one or more pre - adaptation natural language understanding models that are utilized when performing semantic analysis of natural language inputs, and the one or more pre - adaptation natural language understanding models occupy a first amount of local disk space in the first assistant device. In some of those implementations, the corresponding subset stored locally on the first assistant device includes at least one additional natural language understanding model that is added to the one or more pre - adaptation natural language understanding models, occupies a second amount of local disk space in the first assistant device, and includes one or more post - adaptation natural language understanding models, where the second amount is greater than the first amount.
[0126] In some implementations, the corresponding subset stored locally on the first assistant device includes the natural language understanding model of the first device that is utilized in semantic analysis for one or more first classifications, and excludes the natural language understanding model that is utilized in semantic analysis for a second classification. In some of those implementations, the corresponding subset stored locally on the second assistant device includes the natural language understanding model of the second device that is utilized at least in semantic analysis for the second classification.
[0127] In some implementations, the corresponding processing capabilities for each of the heterogeneous assistant devices in an assistant device group include corresponding processor values based on the capabilities of one or more on - device processors, corresponding memory values based on the size of the on - device memory, and / or corresponding disk space values based on the available disk space.
[0128] In some implementations, generating an assistant device group of heterogeneous assistant devices is in response to a user interface input that explicitly indicates a desire to group the heterogeneous assistant devices.
[0129] In some implementations, generating an assistant device group of heterogeneous assistant devices is automatically performed in response to a determination that the heterogeneous assistant devices meet one or more proximity conditions with respect to each other.
[0130] In some implementations, generating an assistant device group of heterogeneous assistant devices is performed in response to receiving a positive user interface input in response to a recommendation to create an assistant device group, the recommendation being automatically generated in response to a determination that the heterogeneous assistant devices meet one or more proximity conditions with respect to each other.
[0131] In some implementations, the method further includes, following assigning a corresponding processing role to each of the heterogeneous assistant devices of the assistant device group, determining that a first assistant device is no longer within the group, and causing the first assistant device to replace a corresponding subset stored locally on the first assistant device with a first set of first on-device models in response to the determination that the first assistant device is no longer within the group.
[0132] In some implementations, determining the collective set is further based on usage data reflecting past usage in one or more of the group's assistant devices.
[0133] In some of those implementations, determining the collective set includes determining a plurality of candidate sets, each of which can be collectively stored locally and collectively used locally by the group's assistant devices, based on corresponding processing capabilities for each of the heterogeneous assistant devices of the assistant device group, and selecting the collective set from the candidate sets based on the usage data.
[0134] In some implementations, a method implemented by one or more processors of an assistant device is provided. The method includes an assistant device determining that it is a lead device for a group of assistant devices in response to a determination that the assistant device is included in a group of assistant devices that includes the assistant device and one or more additional assistant devices. The method further includes, in response to the determination that the assistant device is the lead device for the group, determining a collective set of on-device models for use in collaboratively and locally processing an assistant request directed to any of the heterogeneous assistant devices of the assistant device group, based on the processing capabilities of the assistant device and the processing capabilities received for each of the one or more additional assistant devices, and for each of the on-device models, determining a corresponding designation indicating which of the group of assistant devices will locally store the on-device model. The method further includes, in response to the determination that the assistant device is the lead device for the group, communicating with the one or more additional assistant devices to cause each of the one or more additional assistant devices to locally store any of the on-device models having a corresponding designation for the additional assistant device, locally storing, in the assistant device, the on-device model having a corresponding designation for the assistant device, and assigning one or more corresponding processing roles to each of the group of assistant devices for collaborative local processing of the assistant request directed to the group.
[0135] These and other implementations of the techniques disclosed herein may optionally include one or more of the following features.
[0136] In some implementations, determining that an assistant device is a lead device for a group includes comparing the processing capabilities of the assistant device to the received processing capabilities for each of one or more additional assistant devices, and determining, based on the comparison, that the assistant device is the lead device for the group.
[0137] In some implementations, a group of assistant devices is created in response to a user interface input that explicitly indicates a desire to group assistant devices.
[0138] In some implementations, the method further includes coordinating local processing of an assistant request, in response to a determination that an assistant device is a lead device for a group and receipt of the assistant request at one or more of the group's assistant devices, using the corresponding processing role assigned to the assistant device.
[0139] In some implementations, a method implemented by one or more processors of an assistant device is provided, which includes determining that the assistant device has been removed from a group of heterogeneous assistant devices. This group is a group that included the assistant device and at least one additional assistant device. At the time the assistant device is removed from the group, the assistant device locally stores a set of on-device models, and the set of on-device models was insufficient to fully process utterances directed to the automated assistant locally on the assistant device. The method further includes, in response to a determination that the assistant device has been removed from the group of assistant devices, causing the assistant device to purge one or more of the set of on-device models and retrieve and locally store one or more additional on-device models. Following retrieving and locally storing the one or more additional on-device models, the assistant device, the one or more additional on-device models, and any remainder of the set of on-device models can be utilized locally on the assistant device to fully process utterances directed to the automated assistant.
Description of Signs
[0140] 101B Device Group 101C Device Group 101D Device Group 108 Local Area Network (LAN) 109 Wide Area Network (WAN) 110A First Assistant Device 110B Second Assistant Device 110C Third Assistant Device 110D Fourth Assistant Device 120A Assistant Client 120B Assistant Client 120C Assistant Client 120D Assistant Client 121 Wake Queue Engine 121A1 Wake / Call Engine 121A2 Wake Queue Engine 121B1 Wake / Call Engine 121C1 Wake / Call Engine 121D1 Wake / Call Engine 121D2 Wake Queue Engine 122A1 ASR Engine 122A2 ASR Engine 122A3 ASR Engine 122A4 ASR Engine 122B1 ASR Engine 122B4 ASR Engine 123A1 NLU Engine 123A2 NLU Engine 123B1 NLU Engine 123B2 NLU Engine 124A1 Fulfillment Engine 124B1 Fulfillment Engine 124B2 Fulfillment Engine 125A1 Text-to-Speech (TTS) Engine 125B1 TTS Engine 126A1 Authentication Engine 126B1 Authentication Engine 126C1 Authentication Engine 126D1 Authentication Engine 127A1 Warm Queue Engine 127A2 Warm Queue Engine 127B1 Warm Queue Engine 127B2 Warm Word Engine 127C1 Warm Queue Engine 127D1 Warm Queue Engine 127D2 Warm Queue Engine 128A1 VAD Engine 128B1 VAD Engine 128C1 VAD Engine 128D1 VAD Engine 131A1 On-Device Wake / Call Model 131A2 Wake Queue Model 131B1 On-Device Wake / Call Model 131C1 On-Device Wake / Call Model 131D1 On-Device Wake / Call Model 131D2 Wake Queue Model 132A1 On-Device ASR Model 132A2 On-Device ASR Model 132A3 ASR Model 132A4 ASR Model 132B1 On-Device ASR Model 132B4 ASR Model 133A1 On-Device NLU Model 133A2 On-Device NLU Model 133B1 On-Device NLU Model 133B2 On-Device NLU Model, Fullfillment Model 134A1 On-Device Fullfillment Model 134B1 On-Device Fullfillment Model 135A1 On-Device TTS Model 135B1 On-Device TTS Model 136A1 On-Device Authentication Model 136B1 On-Device Authentication Model 136C1 On-Device Authentication Model 136D1 On-Device Authentication Model 137A1 On-Device Warm Queue Model 137A2 Warm Queue Model 137B1 On-Device Warm Queue Model 137B2 Warm Word Model 137C1 On-Device Warm Queue Model 137D1 On-Device Warm Queue Model 137D2 Warm Queue Model 138A1 On-Device VAD Model 138B1 On-Device VAD Model 138C1 On-Device VAD Model 138D1 On-Device VAD Model 138D2 VAD Model 140 Cloud-Based Automated Assistant Component 150 Local Model Repository 200 Method 300 Method 410 Computing Device 412 Bus Subsystem 414 Processor 416 Network Interface Subsystem 420 User Interface Output Device 422 User Interface Input Device 425 Memory Subsystem. Storage Subsystem 426 File Storage Subsystem 430 Main Random Access Memory (「RAM」) 432 Read Only Memory (「ROM」)
Claims
1. A method implemented by one or more processors, comprising: generating an assistant device group of heterogeneous assistant devices, wherein the heterogeneous assistant devices include at least a first assistant device and a second assistant device, and at the time of generating the group, the first assistant device includes a first set of on-device models locally stored and used when locally processing an assistant request directed to the first assistant device, and the second assistant device includes a second set of on-device models locally stored and used when locally processing an assistant request directed to the second assistant device; determining a collective set of on-device models locally stored for use when cooperatively and locally processing an assistant request directed to any of the heterogeneous assistant devices in the assistant device group, based on corresponding processing capabilities for each of the heterogeneous assistant devices in the assistant device group; in response to the generation of the assistant device group, locally storing, in each of the heterogeneous assistant devices, a corresponding subset of the collective set of on-device models, including purging one or more first on-device models of the first set to provide storage space for the corresponding subset locally stored on the first assistant device, and purging one or more second on-device models of the second set to provide storage space for the corresponding subset locally stored on the second assistant device; assigning one or more corresponding processing roles to each of the heterogeneous assistant devices in the assistant device group, wherein each of the processing roles utilizes one or more corresponding models of the on-device models locally stored. Subsequent to assigning the corresponding processing role to each of the heterogeneous assistant devices of the assistant device group, detecting a vocal utterance via a microphone of at least one device of the heterogeneous assistant devices of the assistant device group; causing, in response to the vocal utterance being detected via the microphone of the assistant device group, the vocal utterance to be locally processed cooperatively by the heterogeneous assistant devices of the assistant device group using their corresponding processing roles A method comprising: **Claim 2** The step of purging, from the first assistant device, one or more first on-device models of the first set includes purging, from the first assistant device, a wake word detection model of a first device of the first set that is used when detecting a first wake word, The corresponding subset stored locally on the second assistant device includes a wake word detection model of a second device that is used when detecting the first wake word, The step of assigning the corresponding processing role includes assigning, to the second assistant device, a first wake word detection role of using the wake word detection model of the second device when monitoring for an occurrence of the first wake word, The vocal utterance includes a first wake word followed by an assistant command, In the first wake word detection role, the second assistant device detects an occurrence of the first wake word and, in response to detecting the occurrence of the first wake word, causes execution of an additional role of the corresponding processing roles. The method according to claim 1. **Claim 3** The additional role of the corresponding processing roles is executed by the first assistant device, The second assistant device causes execution of the additional role of the corresponding processing roles by transmitting an indication of detection of the first wake word to the first assistant device. The method according to claim 2. **Claim 4** The corresponding subset stored locally on the first assistant device includes a first wake word detection model of the first device that is used when detecting one or more first wake words, and excludes a wake word detection model that is used when detecting one or more second wake words. The corresponding subset stored locally on the second assistant device includes a second wake word detection model of the second device that is used when detecting the one or more second wake words, and excludes a wake word detection model that is used when detecting the one or more first wake words. The step of assigning the corresponding processing role includes the step of assigning to the first assistant device a first wake word detection role that uses the wake word detection model of the first device when monitoring for occurrences of the one or more first wake words. The method according to claim 1, wherein the step of assigning the corresponding processing role includes the step of assigning to the second assistant device a second wake word detection role that uses the wake word detection model of the second device when monitoring for occurrences of the one or more second wake words.
5. The corresponding subset stored locally on the first assistant device includes a first language speech recognition model that is used when performing speech recognition in the first language, and excludes a speech recognition model that is used when recognizing speech in the second language. The corresponding subset stored locally on the second assistant device includes a second language speech recognition model that is used when performing speech recognition in the second language, and excludes a speech recognition model that is used when recognizing speech in the first language. The step of assigning the corresponding processing role includes the step of assigning to the first assistant device a first language speech recognition role that uses the first language speech recognition model when performing speech recognition in the first language. The step of assigning the corresponding processing role includes the step of assigning to the second assistant device a second language speech recognition role that utilizes the second language speech recognition model when performing speech recognition in the second language, according to any one of claims 1 to 4.
6. The corresponding subset stored locally on the first assistant device includes a first part of the speech recognition model that is utilized when performing a first part of speech recognition, excluding a second part of the speech recognition model. The corresponding subset stored locally on the second assistant device includes the second part of the speech recognition model that is utilized when performing the second part of speech recognition, excluding the first part of the speech recognition model. The step of assigning the corresponding processing role includes the step of assigning to the first assistant device a first part of a language speech recognition role that utilizes the first part of the speech recognition model when generating a corresponding embedding of the corresponding speech, and the step of transmitting the corresponding embedding to the second assistant device. and The step of assigning the corresponding processing role includes the step of assigning to the second assistant device a second language speech recognition role that utilizes the corresponding embedding from the first assistant device and the second part of the speech recognition model when generating a corresponding recognition of the corresponding speech, according to any one of claims 1 to 4.
7. The corresponding subset stored locally on the first assistant device includes a speech recognition model that is utilized when performing a first part of speech recognition. The step of assigning the corresponding processing role includes the step of assigning to the first assistant device a first part of a language speech recognition role that utilizes the speech recognition model when generating an output, and the step of transmitting the corresponding output to the second assistant device. and The step of assigning the corresponding processing role includes the step of assigning, to the second assistant device, a second language speech recognition role that performs beam search on the corresponding output from the first assistant device when generating the corresponding recognition of the corresponding speech, according to any one of claims 1 to 4.
8. The corresponding subset stored locally on the first assistant device includes one or more pre-adapted natural language understanding models that are utilized when performing semantic analysis of natural language input. The one or more pre-adapted natural language understanding models occupy a first amount of local disk space in the first assistant device. The corresponding subset stored locally on the first assistant device includes one or more post-adapted natural language understanding models that include at least one additional natural language understanding model added to the one or more pre-adapted natural language understanding models. The one or more post-adapted natural language understanding models occupy a second amount of the local disk space in the first assistant device, and the second amount is greater than the first amount, according to any one of claims 1 to 7.
9. The corresponding subset stored locally on the first assistant device includes the natural language understanding model of the first device that is utilized in semantic analysis for one or more first classifications, and excludes the natural language understanding model that is utilized in semantic analysis for a second classification. The corresponding subset stored locally on the second assistant device includes the natural language understanding model of the second device that is utilized at least in semantic analysis for the second classification, according to any one of claims 1 to 7.
10. The corresponding processing capabilities for each of the heterogeneous assistant devices in the assistant device group include corresponding processor values based on the capabilities of one or more on-device processors, corresponding memory values based on the size of the on-device memory, and corresponding disk space values based on the available disk space, according to any one of claims 1 to 9.
11. The step of generating the assistant device group of heterogeneous assistant devices responds to a user interface input that explicitly indicates a request to group the heterogeneous assistant devices, the method according to any one of claims 1 to 10.
12. The step of generating the assistant device group of heterogeneous assistant devices is automatically executed in response to a determination that the heterogeneous assistant devices satisfy one or more proximity conditions with respect to each other, the method according to any one of claims 1 to 10.
13. Subsequent to assigning the corresponding processing role to each of the heterogeneous assistant devices of the assistant device group, Determining that the first assistant device is no longer within the group; In response to the determination that the first assistant device is no longer within the group, Causing the first assistant device to replace the corresponding subset stored locally on the first assistant device with the first on-device model of the first set The method according to any one of claims 1 to 10, further comprising.
14. The step of determining the collective set is further based on usage data reflecting past usage in one or more of the assistant devices of the group, the method according to any one of claims 1 to 13.
15. The step of determining the collective set is Based on the corresponding processing capabilities for each of the heterogeneous assistant devices of the assistant device group, determining a plurality of candidate sets that can each be collectively stored locally and collectively used locally by the assistant devices of the group; Selecting the collective set from the candidate sets based on the usage data The method according to claim 14, comprising.
16. A method implemented by one or more processors of an assistant device, comprising: In response to a determination that the assistant device is included within a group of assistant devices including the assistant device and one or more additional assistant devices, determining that the assistant device is a lead device for the group; in response to determining that the assistant device is the lead device for the group, based on the processing capabilities of the assistant device and the received processing capabilities for each of the one or more additional assistant devices, a collective set of on-device models for use in cooperatively and locally processing an assistant request directed to any of the heterogeneous assistant devices of the group of assistant devices; for each of the on-device models, a corresponding designation indicating which of the assistant devices of the group will store the on-device model locally; determining; communicating with the one or more additional assistant devices to cause each of the one or more additional assistant devices to locally store any of the on-device models having the corresponding designation for the additional assistant device; locally storing, in the assistant device, the on-device model having the corresponding designation for the assistant device; assigning one or more corresponding processing roles to each of the assistant devices of the group for cooperative local processing of an assistant request directed to the group; a method comprising.
17. The step of determining that the assistant device is the lead device for the group comprises: comparing the processing capabilities of the assistant device with the received processing capabilities for each of the one or more additional assistant devices; determining, based on the comparison, that the assistant device is the lead device for the group; The method according to claim 16, comprising.
18. The method according to claim 16 or claim 17, wherein the group of assistant devices is created in response to a user interface input explicitly indicating a requirement to group the assistant devices.
19. In response to a determination that the assistant device is the lead device for the group and receipt of an assistant request at one or more of the assistant devices of the group, adjusting coordinated local processing of the assistant request using the corresponding processing role assigned to the assistant device The method according to any one of claims 16 to 18, further comprising.
20. A method implemented by one or more processors of an assistant device, comprising: determining that the assistant device has been removed from a group of heterogeneous assistant devices, the group already includes the assistant device and at least one additional assistant device, at the time when the assistant device is removed from the group, the assistant device locally stores a set of on-device models, the set of on-device models is insufficient to fully process the utterances directed to the automated assistant locally at the assistant device, the step of in response to a determination that the assistant device has been removed from the group of assistant devices, purging one or more of the set of on-device models from the assistant device and retrieving and locally storing one or more additional on-device models, subsequent to retrieving and locally storing the one or more additional on-device models, the assistant device, the one or more additional on-device models, and any remainder of the set of on-device models can be utilized locally at the assistant device to fully process the utterances directed to the automated assistant, the step of including a method.
Citation Information
Patent Citations
Picture processor in modular-type picture processing system, recording medium which records memory management program applied to processor and which computer can read and memory management method
JP2000132400A
Speaker matching using collocation information
JP2017517027A
Information processing system, information processing device, information processing method and program
JP2018173515A