Unsupervised Federated Learning of Machine Learning Model Layers

Through unsupervised federated learning and gradient training of public data, richer and more robust feature representations are generated, which solves the problem of poor model performance in the prior art, and trains the combined model through supervised learning, improving the accuracy and efficiency of the model.

CN116134453BActive Publication Date: 2025-06-17GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202080104721.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-07-20
Publication Date
2025-06-17
Estimated Expiration
2040-07-20

AI Technical Summary

Technical Problem

In federated learning, existing machine learning models have poor performance due to pre-training using agents or biased data, and cannot effectively reflect data characteristics in the deployment environment.

Method used

Unsupervised or self-supervised federated learning methods are used to generate gradients locally on client devices and train in remote systems. At the same time, additional gradients are generated using publicly available data to further train the global machine learning model layer. The trained global model layer is then combined with the additional layer to generate a combined machine learning model, and further trained using supervised learning.

Benefits of technology

Through unsupervised federated learning, the global model layer is trained, combined with gradient training of public data, and a richer and more robust feature representation is generated, which improves the generalization ability and performance of the model. At the same time, supervised learning is used to train the combined model, which improves the accuracy and efficiency of the model on specific tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116134453B_ABST
    Figure CN116134453B_ABST
Patent Text Reader

Abstract

The embodiments disclosed herein are directed to unsupervised federated training of a global machine learning (“ML”) model layer, which can be combined with additional layers after federated training to produce a combined ML model. A processor can: detect audio data that captures an oral utterance of a user of a client device; process the audio data using a local ML model to generate a prediction output; generate gradients based on the prediction output using unsupervised learning local to the client device; transmit the gradients to a remote system; update weights of the global ML model layer based on the gradients; after updating the weights, remotely train, on the remote system, the combined ML model, the combined ML model including the updated global ML model layer and additional layers; transmit the combined ML model to the client device; and make predictions on the client device using the combined ML model.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Federated learning of machine learning models is an increasingly popular machine learning technique for training machine learning models. In traditional federated learning, local machine learning models are stored locally on a user's client device, while a global machine learning model, which is a cloud-based counterpart of the local machine learning model, is stored remotely in a remote system (e.g., a server cluster). Using the local machine learning model, the client device can process user input detected at the client device to generate a predicted output, and can compare the predicted output with the ground truth output to generate gradients using supervised learning techniques. Additionally, the client device can transmit the gradients to the remote system. The remote system can utilize the gradients and optionally additional gradients generated in a similar manner at additional client devices to update the weights of the global machine learning model. Further, the remote system can transmit the global machine learning model or the updated weights of the global machine learning model to the client device. The client device can then replace the local machine learning model with the global machine learning model, or replace the weights of the local machine learning model with the updated weights of the global machine learning model, thereby updating the local machine learning model. The local and global machine learning models typically each include a feature extractor portion combined with additional layers. The combined model can optionally be pre-trained using surrogate data before being used in federated learning. Any pre-training can be performed on the remote system (or an additional remote system) and uses supervised learning without using any gradients generated by the client device. This pre-training is generally based on surrogate or biased data, which may not reflect the data that will be encountered when deploying the machine learning model, resulting in poor performance of the machine learning model. Summary of the Invention

[0002] Some embodiments disclosed herein are directed to unsupervised (or self-supervised) federated learning of a machine learning (“ML”) model layer. The ML model layer can be trained in a remote system based on gradients generated locally at client devices using unsupervised (or self-supervised) learning techniques and transmitted to the remote system. The ML model layer can be optionally further trained in the remote system based on additional gradients generated remotely in the remote system using unsupervised (or self-supervised) learning techniques and based on publicly available data. Some embodiments disclosed herein additionally or alternatively are directed to combining the ML model layer with an additional upstream layer after training the ML model layer, thereby producing a combined machine learning model. Some of these embodiments are further directed to training the combined machine learning model (e.g., at least its additional upstream layer) at the remote system using supervised learning techniques. Thus, various embodiments disclosed herein seek to first train the ML model layer using unsupervised (or self-supervised) federated learning, combine the trained ML model layer with an additional upstream layer to generate a combined model, and then train the combined model using supervised learning (e.g., non-federated supervised learning performed on the remote system). It is noted that this is in contrast to alternative techniques that use supervised learning and / or do not use federated learning to pre-train the upstream layer, combine the pre-trained ML model layer with additional layers to generate a combined model, and then only use federated learning in training the combined model.

[0003] In some embodiments, the ML model layer is used to process audio data. For example, the ML model layer can be used to process the Mel filter bank features of the audio data and / or other representations of the audio data. In some of those embodiments, when generating gradients at the client device, the client device can: detect, via a corresponding microphone, audio data capturing the spoken words of the corresponding user of the client device; process the audio data (e.g., its features) using a local machine learning model, the local machine learning model including an ML model layer corresponding to the global ML model layer and used to generate an encoding of the audio data, to generate a prediction output; and generate a gradient based on the prediction output using unsupervised (or self-supervised) learning. The local machine learning model can include, for example, an encoder-decoder network model (“encoder-decoder model”), a deep belief network (“DBN model”), a generative adversarial network model (“GAN model”), a cycle generative adversarial network model (“CycleGAN model”), a transformer model, a prediction model, other machine learning models including an ML model layer for generating an encoding of the audio data, and / or any combination thereof. Additionally, the ML model layer for generating an encoding of the audio data can be part of one of these models (e.g., all or part of the encoder in an encoder-decoder model), and as described above, corresponds to the global ML model layer (e.g., having the same structure as the global ML model layer). As a non-limiting example, if the local machine learning model is an encoder-decoder model, the part for generating an encoding of the audio data can be the encoder part or a downstream layer of the encoder-decoder model. As another non-limiting example, if the local machine learning model is a GAN model or a CycleGAN model, the part for generating an encoding of the audio data can be the real-to-encoding generator model of the GAN model or the CycleGAN model. Although the previous examples are described for an ML model layer for processing data to generate an encoding of that data, the ML model layer can alternatively be used to process data to generate an output that is the same (or even higher-dimensional than the processed data) as the processed data. For example, the input layer of the ML model layer can have a specific dimension, and the output layer of the ML model can also have the same specific dimension. As a specific example, the ML model layer can be one or more downstream layers of a local machine learning model, and the output generated using those ML model layers can be provided to the upstream encoder-decoder part of the local machine learning model.

[0004] In some versions of those embodiments, the generation of the gradient can be based on the predicted output generated across local machine learning models given the audio data of the corresponding user's spoken utterance of a capture client device. As a non-limiting example, assume that the local machine learning model is an encoder-decoder model and the encoder part is used to generate the encoding of the audio data, while the decoder part is used to process the encoding of the audio data to attempt to reconstruct the audio data. In this example, the predicted output can be generated by using the decoder part to process the encoding of the audio data and seeking the predicted audio data corresponding to the audio data based on which the encoding was generated. In other words, the predicted output generated using the encoder-decoder model can be the predicted audio data generated using the decoder and seeking to correspond to the spoken utterance. Thus, the client device can determine the difference between the audio data and the predicted audio data. For example, the client device can determine the difference between the audio data and the predicted audio data, such as based on comparing the differences in their analog audio waveforms, the differences between the representations of the audio data and the predicted audio data in the latent space, the differences based on comparing the features of the audio data and the predicted audio data determined by deterministic calculations (e.g., their mel filter bank features, their Fourier transforms, their mel cepstral frequency coefficients, and / or other representations of the audio data and the predicted audio data). In some embodiments of determining the difference between the representations of the audio data and the predicted audio data in the latent space, more useful features can be extracted therefrom compared to determining the difference based on comparing the original audio data. Additionally, the client device can generate a gradient based on the determined difference between the audio data and the predicted audio data.

[0005] As another non-limiting example, assume that the spoken utterance includes a first part of the audio data and a second part of the audio data following the first part. Further assume that the local machine learning model is an encoder prediction model, the encoder part is used to generate the encoding of the first part, and the prediction part is used to process the encoding in generating the predicted second part of the audio data. In this example, the predicted output can be the predicted second part of the audio data, which seeks to correspond to the actual second part of the audio data of the corresponding user's spoken utterance of the client device. In other words, the encoder prediction model can be used to process the first part of the audio data and generate the predicted second part of the audio data, which seeks to correspond to the actual second part of the audio data. Thus, the client device can determine the difference between the actual second part of the audio data and the predicted second part of the audio data in a similar manner as described in the previous example. Additionally, the client device can generate a gradient based on the determined difference between the actual second part of the audio data and the predicted second part of the audio data. It is noted that each client device can generate the corresponding gradient in this way and can transmit the corresponding gradient to the remote system.

[0006] In some versions of those embodiments, the transmission of gradients from the client device to the remote system is in response to the client device determining that one or more conditions are met. The one or more conditions can include, for example, that the client device has authorized the transmission of gradients, the client device is charging, the client device has at least a threshold charge state, the temperature of the client device (based on temperature sensors on one or more devices) is below a threshold, the client device is not being held by a user, a time condition associated with the client device (e.g., between specific time periods, every N hours (where N is a positive integer), etc.) and / or other time conditions associated with the client device, whether a given one of the client devices has generated a threshold number of gradients, and / or other conditions. For example, in response to a given one of the client devices determining that it is in a charging state and in response to a given one of the client devices determining that the current time at the location of the given one of the client devices is between 2:00 am and 5:00 am, then the given one of the client devices can transmit the corresponding gradient to the remote system. As another example, in response to a given one of the client devices determining that one or more corresponding gradients have been generated and in response to a given one of the client devices determining that it last transmitted gradients to the remote system 7 days ago, then the given one of the client devices can transmit one or more corresponding gradients to the remote system, assuming that the given one of the client devices has authorized the transmission of gradients. In various embodiments, the client device determines whether one or more conditions are met in response to a request for gradients from the remote system. For example, if a given one of the client devices determines that one or more conditions are met in response to receiving a request for gradients from the remote system, then the given one of the client devices can transmit the gradients. However, if one or more conditions are not met, then the given one of the client devices can avoid transmitting the gradients to the remote system.

[0007] In some of those embodiments, the remote system can receive gradients from the client device and update one or more weights of the global ML model layer based on the gradients, thereby producing an updated global ML model layer. As described above, the remote system can also optionally generate additional gradients (i.e., remote gradients) using unsupervised (or self-supervised) learning techniques based on publicly available data and optionally utilize the additional gradients to update one or more weights of the global ML model layer. In some versions of those embodiments, the additional gradients can be generated in a similar manner as described above with respect to the client device. However, instead of generating gradients based on processing audio data of captured dictated speech on the client device using a local machine learning model, the remote system can generate the additional gradients by processing additional audio data captured in publicly available resources (e.g., from video sharing platforms, audio sharing platforms, image sharing platforms, and / or any non-access-restricted publicly available resources) using the global machine learning model (e.g., the global version of the local machine learning model). For the dictated speech of the corresponding user, the global ML model layer is updated using different data, resulting in a more robust global ML model layer than if the global ML model layer were updated only based on the dictated speech of the corresponding user. For example, the global ML model layer can be updated to generate a richer representation of speech features because it is not overly biased towards the dictated speech typically received on the client device.

[0008] In some further versions of those embodiments, the remote system can assign the gradients and / or additional gradients to specific iterations of the update of the global ML model layer. The remote system can assign the gradients and / or additional gradients to specific iterations based on one or more criteria. The one or more criteria can include, for example, a threshold number of gradients and / or additional gradients, a threshold duration of updating using the gradients and / or additional gradients, and / or other criteria. In other versions of these embodiments, the remote system can assign the gradients to individual subsets of specific iterations. Each subset can optionally include gradients from at least one unique client device not included in another subset. As a non-limiting example, the remote system can update the global ML model layer using the following: 100 gradients received from a first subset of client devices, then 100 gradients generated at the remote system based on publicly available resources, then another 100 gradients received from a second subset of client devices, and so on. As another non-limiting example, the remote system can update the global ML model layer by the following: updating for one hour based on gradients received from client devices, then updating for one hour based on gradients generated at the remote system based on publicly available resources, then updating for another hour based on gradients received from client devices, and so on. It is noted that these threshold numbers and / or durations can vary between the gradients received from client devices and the additional gradients generated at the remote system.

[0009] In some versions of those embodiments, after training the global ML model layer, the remote system can combine the global ML model layer with additional layers to generate a combined machine learning model. The additional layers can correspond to downstream layers of, for example, a voice activity detection model, a hotword detection model, a speech recognition model, a continuous conversation model, a no-hotword detection model, a gaze detection model, a lip movement detection model, an object detection model, an object classification model, a face recognition model, and / or other machine learning models. Notably, the same global ML model layer can be combined with additional layers of multiple different types of machine learning models, resulting in multiple different combined machine learning models that use the same global ML model layer. It should also be noted that the additional layers of the combined machine learning model may be structurally different from those of the local machine learning models used to generate gradients for the federated learning of the global ML model layer. For example, the local machine learning model can be an encoder-decoder model that has an encoder portion that is structurally corresponding to the global ML model layer. The combined machine learning model includes the global ML model layer that is structurally corresponding to the encoder portion, but the additional layers of the combined machine learning model may be structurally different from the decoder portion of the local machine learning model. For example, the additional layers can include more or fewer layers than the decoder portion, different connections between the layers, different output dimensions, and / or different types of layers (e.g., recurrent layers instead of feedforward layers).

[0010] In addition, the remote system can use supervised learning techniques to train the combined machine learning model. For example, the identified supervised training instances (i.e., with labeled ground truth outputs) can be used to train a given combined machine learning model. Each supervised training instance can be identified based on the ultimate goal of training the given combined machine learning model. As a non-limiting example, assuming that the given combined machine learning model is trained as a speech recognition model, the identified supervised training instances can correspond to the training instances used to train the speech recognition model. Thus, in this example, the global ML model layer of the speech recognition model can process the training instance input (e.g., audio data corresponding to speech) to generate a feature representation corresponding to the training instance input, and the additional layers of the given speech recognition model can process the feature representation corresponding to the speech to generate a predicted output (e.g., predicted phonemes, predicted tokens, and / or recognized text corresponding to the speech of the training instance input). The predicted output can be compared with the training instance output (e.g., ground truth phonemes, tokens, and / or text corresponding to the speech), and can be backpropagated across the speech recognition model including the global ML model layer and the additional layers of the given speech recognition model, and / or used to update the weights of only the additional layers of the combined machine learning model. This process can be repeated using multiple training instances to train the speech recognition model. In addition, this process can be repeated using the corresponding training instances to generate other combined machine learning models (e.g., for voice activity detection, hot word detection, continued conversation, hot wordless call, gaze detection, mouth movement detection, object detection model, object classification model, face recognition, and / or other models).

[0011] In some versions of those embodiments, once the combined machine learning model is trained in the remote system, the remote system can transmit the trained combined machine learning model back to the client device. The client device can utilize the trained machine learning model to make predictions based on user input detected at the client device of the corresponding user receiving the combined machine learning. The predictions can be based on an additional layer combined with the updated global ML model layer transmitted to the client device. As a non-limiting example, assume that the additional layers are those of a speech recognition model and the combined machine learning model is trained using speech recognition training instances. In this example, the user input on the client device can be an oral utterance, and the prediction on the client device can be the predicted phonemes generated by processing the oral utterance using the combined machine learning model. As another non-limiting example, assume that the additional layers are those of a hotword detection model and the combined machine learning model is trained using hotword detection training instances. In this example, the user input on the client device can be an oral utterance including a hotword, and the prediction made on the client device can include an indication of whether the oral utterance includes the hotword generated by processing the oral utterance using the combined machine learning model. In some further versions of these embodiments, the remote system can also transmit the updated global ML model layer to the client device to replace the portion of the local machine learning model used to generate the audio data encoding.

[0012] In various embodiments, the global ML model layer can be continuously updated based on gradients generated using the unsupervised (or self-supervised) learning techniques described herein, resulting in an updated global ML model layer. Additionally, the global ML model layer can be combined with additional layers of other machine learning models described herein, resulting in an updated combined machine learning model. Additionally, the updated combined machine learning model can be trained using the supervised learning techniques described herein. The updated combined machine learning model can then be transmitted to the client device and stored locally on the client device (and optionally replace the corresponding one in the combined machine learning model at the client device if it exists) for subsequent use by the client device. In some additional and / or alternative embodiments, the existing global ML model layer of a combined machine learning model can be replaced with the updated global ML model layer, resulting in a modified combined machine learning model. The modified combined machine learning model can then be transmitted to the client device to replace the corresponding one of the combined machine learning model at the client device.

[0013] By first training the global ML model layer in this manner and then combining it with additional layers of other machine learning models, the resulting combined machine learning model can be trained more effectively. For example, as described herein, using unsupervised (or self-supervised) federated learning to train the global ML model layer can result in a global ML model layer that can be used to generate rich encodings for various oral discourse features. For example, when a feature extractor model is used to process audio data, training the global ML model layer using gradients from a variety of different client devices results in gradient-based training based on audio data of different voices of different users, audio data with different background noise conditions, and / or audio data generated by different client device microphones (which may have different properties and result in different acoustic features), etc. This may result in the global ML model layer, once trained, being usable to generate richer and / or more robust encodings for multiple different applications (e.g., speech recognition, hotword detection, dictation, voice activity detection, etc.). Additionally, embodiments of additional training based on gradients from publicly available audio data can prevent the feature extractor from being overly biased towards certain terms and / or certain speech styles. This may also result in the global ML model layer, once trained, being usable to generate richer and / or more robust encodings.

[0014] Such rich and / or robust encodings can enable the resulting combined machine learning model to converge faster during training. Such rich encodings can additionally and / or alternatively enable the resulting combined machine learning model to achieve high recall and / or high precision using a smaller number of training instances. In particular, when the updated global ML model layer is combined with additional layers, since the updated global ML model layer is updated based on unlabeled data without any direct supervision, the resulting combined machine learning model can achieve high recall and / or high precision using a smaller number of labeled training instances. Additionally, by training the global ML model layer using unsupervised (or self-supervised) learning, the training of the global ML model layer can be computationally more efficient in the sense that it does not require labeled training instances. Additionally, the need for any on-device supervision from user corrections or user actions, which may be difficult or impossible to obtain, is avoided.

[0015] Although the various embodiments above are described with respect to processing audio data and / or speech processing, it should be understood that this is for illustrative purposes and not for limitation. As described in detail herein, the techniques described herein can additionally or alternatively be applied to a global ML model layer, which can be used to process additional and / or alternative types of data and generate corresponding encodings. For example, techniques for image processing can be utilized so that the global ML model layer can be trained to generate encodings of images. For example, assume that the local machine learning model is an encoder-decoder model, which includes: an encoder portion for generating an encoding based on image data generated at a given one of the client devices; and, a decoder portion for reconstructing an image based on the encoding of the image data, thereby obtaining predicted image data corresponding to the image. A given one of the client devices can compare the image data with the predicted image data to generate gradients using unsupervised (or self-supervised) learning and transmit the gradients to the remote system. The remote system can update one or more weights of the global ML model layer based on the gradients (and additional gradients and / or remote gradients from the client devices), and can combine the global ML model layer with additional layers of the given image processing machine learning model, thereby producing a combined machine learning model for image processing. The remote system can further train the combined machine learning model using supervised learning and transmit the combined machine learning model for image processing to the client device. For example, the combined machine learning model can be the one that is trained to predict the location of an object in an image, the classification of an object in an image, the captioning of an image, and / or other predictions.

[0016] Accordingly, the above description is provided as an overview of some embodiments of the present disclosure. Further descriptions of those embodiments and other embodiments are described in more detail below. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1A 、 1B Figures 1A, 1B, 1C, and 1D depict example process flows showing various aspects of the present disclosure in accordance with various embodiments.

[0018] Figure 2 depicts a block diagram of an example environment including various components from Figure 1A 、 1B Figures 1A, 1B, 1C, and 1D and in which embodiments disclosed herein can be implemented.

[0019] Figure 3 describes a flowchart that illustrates an example method in accordance with various embodiments: generating gradients locally at a client device using unsupervised learning; transmitting the generated gradients to a remote system, which utilizes the generated gradients to update weights of a global machine learning model layer; and receiving at the client device a combined machine learning model including the updated global machine learning model layer and additional layers.

[0020] Figure 4 depicts a flow chart that illustrates an example method according to various embodiments: updating weights of a global machine learning model layer based on gradients received from multiple client devices and / or generated at a remote system based on publicly available data; generating a combined machine learning model that includes the updated global machine learning model layer and additional layers; and transmitting the combined machine learning model to one or more of the multiple client devices.

[0021] Figure 5 depicts a flow chart that illustrates an example method according to various embodiments: generating gradients at a client device using unsupervised learning; updating weights of a global machine learning model layer based on the gradients; training a combined machine learning model that includes the updated global machine learning model layer and additional layers; and using the combined machine learning model at the client device to make predictions based on user input detected at the client device.

[0022] Figure 6 depicts an example architecture of a computing device according to various embodiments. Detailed Description

[0023] Figure 1A-1D depicts an example process flow showing various aspects of the present disclosure. Figure 1A illustrates client device 110, and client device 110 includes components contained within the box in Figure 1A which represents client device 110. The local machine learning engine 122 may detect audio data 101 corresponding to oral discourse detected (or stored in the oral discourse database 101N) via one or more microphones of client device 110 and / or may detect image data 102 corresponding to non-verbal physical movements (e.g., gestures and / or motions, body postures and / or body movements, eye gazes, facial movements, mouth movements, etc.) detected (or stored in the image database 102N) via one or more non-microphone sensor components of client device 110. One or more non-microphone sensors may include vision components, proximity sensors, pressure sensors, and / or other sensors capable of generating image data 102. The local machine learning engine 122 processes the audio data 101 and / or the image data 102 using the local machine learning model 152A to generate a prediction output 103.

[0024] The local machine learning model 152A may include, for example, an encoder-decoder network model (“encoder-decoder model”), a deep belief network (“DBN model”), a generative adversarial network model (“GAN model”), a cyclic generative adversarial network model (“CycleGAN model”), a transformer model, a prediction model, and other machine learning models including machine learning model layers for processing Mel filter bank features of the audio data 101, other representations of the audio data 101, representations of the image data 102, and / or any combination thereof used by the local machine learning engine 122. Notably, the machine learning model layers of the local machine learning model 152A can be used to generate encodings of the audio data 101 and / or the image data 102, or fed to downstream layers that generate encodings of the audio data. Additionally, note that this portion structurally corresponds to the global machine learning model 152B (described below). Further, an additional portion of the local machine learning model is used to generate a prediction output 103 based on the encodings of the audio data 101 and / or the image data 102. For example, if a given local machine learning model utilized by the local machine learning engine 122 is an encoder-decoder model, then the portion used to generate encodings of the audio data 101 and / or the image data 102 can be the encoder portion or a downstream layer of the encoder-decoder model, and the additional portion used to generate the prediction output 103 can be the decoder portion of the encoder-decoder model. As another example, if a given local machine learning model utilized by the local machine learning engine 122 is a CycleGAN model, then the portion used to generate encodings of the audio data 101 and / or the image data 102 can be the real-to-encoding generator portion of the CycleGAN model, and the additional portion used to generate the prediction output 103 can be the encoding-to-real generator portion of the CycleGAN model.

[0025] In some embodiments, when the local machine learning engine 122 generates the prediction output 103, it can be locally stored on the client device in association with the corresponding audio data 101 and / or image data 102 in an on-device storage (not depicted) for a threshold amount of time (e.g., days, weeks, months, etc.). This allows the corresponding audio data and / or image data 102 to be processed at various times on the client device 110 within the threshold amount of time to generate the corresponding prediction output. In some versions of those embodiments, the prediction output 103 can be retrieved later (such as when one or more conditions described herein are met) by the unsupervised learning engine 126 for use in generating the gradient 104. The unsupervised learning engine 126 can generate the gradient 104 locally on the client device 110 using unsupervised learning techniques, as described in more detail herein (e.g., with respect to Figure 1B and 1C)。These unsupervised learning techniques can also be considered "self-supervised" learning techniques because the unsupervised learning engine 126 learns to extract certain features from the audio data 101, the image data 102, and / or other data generated locally on the client device 110. Additionally, the storage on the device can include, for example, read-only memory (ROM) and / or random access memory (RAM) (e.g., as Figure 6 shown). In other embodiments, the prediction output 103 can be provided to the unsupervised learning engine 126 in real time.

[0026] In various embodiments, the generation of the gradient 104 can be based on the prediction output 103 generated across the local machine learning engine 122 given the audio data 101 and / or the image data 102. In some embodiments, and as Figure 1B shown, the prediction output 103 can be the predicted audio data 103A. In some versions of those embodiments, the local machine learning engine 122 can process the audio data 101 using the encoding engine 122A to generate an encoding of the audio data 201. Additionally, the local machine learning engine 122 can process the encoding of the audio data 201 using the decoding engine 122B to generate the predicted audio data 103A. The encoding engine 122A and the decoding engine 122B can be part of one or more machine learning models stored in the local machine learning model database 152A1. Additionally, the encoding of the audio data 201 can be a feature representation of the audio data, such as, for example, a tensor of values, such as a vector or matrix of real numbers, optionally in a reduced-dimensional space where the dimensions are reduced relative to the dimensions of the audio data. As a non-limiting example, the encoding of the audio data 201 can be a vector of 128 values, such as values that are each real numbers from 0 to 1. As another non-limiting example, the encoding of the audio data can be a 2x64 value matrix or a 3x64 value matrix. For example, the encoding engine 122A can utilize the encoder portion of an encoder-decoder model to generate the encoding of the audio data 201, and the decoding engine 122B can utilize the decoder portion of the encoder-decoder model to generate the predicted audio data 103A. As another example, the encoding engine 122A can utilize the real-to-encoding generator portion of a CycleGAN model to generate the encoding of the audio data 201, and the decoding engine 122B can utilize the encoding-to-real generator portion of the CycleGAN model to generate the predicted audio data 103A.

[0027] Notably, as Figure 1BAs shown, the processing of encoding the prediction audio data 103A based on the audio data 201 seeks to correspond to the audio data 101. In other words, the encoding engine 122A generates an intermediate representation of the audio data 101 (e.g., the encoding of the audio data 201), and the decoding engine 122B seeks to reconstruct the audio data 101 from this intermediate representation when generating the prediction audio data 103A. Therefore, assuming that there are no errors in the local machine learning engine 122 during encoding and decoding, there should be little change between the audio data 101 and the prediction audio data 103A.

[0028] Then, the unsupervised learning engine 126 can use unsupervised learning to process the audio data 101 and the prediction audio data 103A to generate the gradient 104. It is worth noting that when using unsupervised learning, there is no labeled data (i.e., no ground truth output) to compare with the audio data 101. Instead, the unsupervised learning engine 126 can utilize the comparison engine 126A to directly compare the audio data 101 and the prediction audio data 103A to determine the difference between them, and the unsupervised learning engine 126 can generate the gradient 104 based on the difference. For example, the comparison engine 126A can compare the audio waveform corresponding to the audio data 101 and the predicted audio waveform corresponding to the prediction audio data 103A to determine the difference between the audio data 101 and the prediction audio data 103A, and the unsupervised learning engine 126 can generate the gradient 104 based on the difference. As another example, the comparison engine 126A can compare the features of the audio data 101 and the deterministically calculated prediction audio data 103A (e.g., its mel filter bank features, its Fourier transform, its mel cepstral frequency coefficients, and / or other representations of the audio data and the prediction audio data) to determine the difference between them, and the unsupervised learning engine 126 can generate the gradient 104 based on this difference. Therefore, the comparison engine 126A can utilize any technique to compare the audio data 101 with the prediction audio data 103A to determine the difference between the two, and the unsupervised learning engine can generate the gradient 104 based on the difference.

[0029] Although the generation of the gradient 104 based on the audio data 101 is described herein Figure 1B, but it should be understood that this is for illustration purposes and does not imply a limitation. As a non - restrictive example, the gradient 104 can be generated additionally or alternatively based on the image data 102. For example, the encoding engine 122A can process the image data 102 to generate an encoding of the image data, and the decoding engine 122B can process the encoding of the image data in a similar manner as described above to generate predicted image data. The encoding of the image data can be a feature representation of the image data, such as, for example, a tensor of values, such as a vector or matrix of real numbers, which is optionally in a reduced - dimensional space where its dimensions are reduced relative to the dimensions of the image data. As a non - restrictive example, the encoding of the image data can be a vector with 128 values, such as values that are each real numbers from 0 to 1. As another non - restrictive example, the encoding of the image data can be a 2x64 matrix of values or a 3x64 matrix of values. The decoding engine 122B can process the encoding of the image data to generate predicted image data. The comparison engine 126A can then compare the image data 102 and the predicted image data or its features to determine the difference between them, and the unsupervised learning engine 126 can generate the gradient 104 based on the difference. In embodiments where the gradient is generated based on the audio data 101 and the image data 102, the gradient can be indexed based on whether it corresponds to the audio data 101 and / or the image 102. This allows updating the weights of various global machine - learning model layers based on one of the audio - based gradients or the image - based gradients, or updating the weights of a single global machine - learning model layer based on both the audio - based gradient and the image - based gradient.

[0030] In some additional and / or alternative embodiments, and as Figure 1C shown, the predicted output 103 can be a predicted second part 103B of the audio data. Notably, as Figure 1CAs shown, the audio data 101 is segmented into a first portion 101A of the audio data and a second portion 101B of the audio data that temporally follows the first portion 101A of the audio data. In some versions of those embodiments, the first portion 101A of the audio data may be the first portion of a spoken utterance (e.g., “What's the weather...”) and the second portion 101B of the audio data may be the second portion of the same spoken utterance that immediately follows the first portion 101B of the audio data (e.g., “...inLouisville, KY”), while in other versions of those embodiments, the first portion 101A of the audio data may be the first spoken utterance (e.g., “What's the weather in Louisville, KY”) and the second portion 101B of the audio data may be the second spoken utterance that immediately follows the first spoken utterance (e.g., “How about in Lexington, KY”).

[0031] In some versions of those embodiments, the local machine learning engine 122 may use the encoding engine 122A to process the first portion 101A of the audio data to generate a signal in accordance with the embodiment of the present invention. Figure 1B The encoding of the first portion of the audio data 202 is generated in a similar manner as described above. Figure 1B In contrast, the local machine learning engine 122 may use the prediction engine 122C instead of the decoding engine 122B to process the encoding of the first portion of the audio data 202 to generate the predicted second portion 103B of the audio data. The prediction engine 122C may utilize one or more prediction models also stored in the local machine learning model database 152A1 to generate the predicted second portion 103B of the audio data. Similar to Figure 1BEncoding of the audio data 201, the encoding of the first part of the audio data 202 can be a feature representation of the audio data, such as, for example, a tensor of values, such as a vector or matrix of real numbers in a reduced-dimensional space where optionally its dimensions are reduced with respect to the dimensions of the audio data. As a non-limiting example, the encoding of the audio data 201 can be a vector of 128 values, such as values each being a real number from 0 to 1. As another non-limiting example, the encoding of the audio data can be a 2x64 value matrix or a 3x64 value matrix. Additionally, the prediction engine 122C can select a given one of one or more prediction models based on the encoding of the first part of the audio data 202. For example, if the encoding of the first part of the audio data 202 includes one or more tokens corresponding to the first part 101A of the audio data, the prediction engine 122C can utilize the first prediction model to generate the predicted second part 103B of the audio data. Conversely, if the encoding of the first part of the audio data 202 includes one or more phonemes corresponding to the first part 101A of the audio data, the prediction engine 122C can utilize a different second prediction model to generate the predicted second part 103B of the audio data. The comparison engine 126A can compare the second part 101B of the audio data and the predicted second part 103B of the audio data in the same manner as described above with respect to Figure 1B to determine the difference between them, and the unsupervised learning engine 126 can generate a gradient 104 based on this difference.

[0032] Although this document also describes Figure 1C generating the gradient 104 based on the audio data 101, it should be understood that this is for illustration purposes and does not imply a limitation. As a non-limiting example, the gradient 104 can be generated based on the image data 102. For example, the encoding engine 122A can process the first part of the image data 102 (or the first image in an image stream) to generate an encoding of the first part of the image data, and the prediction engine 122C can process the encoding of the first part of the image data in a similar manner as described above to generate the predicted second part of the image data. The comparison engine 126A can compare the second part 102 of the image data and the predicted second part of the image data to determine the difference between them, and the unsupervised learning engine 126 can generate a gradient 104 based on the difference. Similar to Figure 1B in an implementation where the gradient is generated based on the audio data 101 and the image data 102 in the manner described with respect to Figure 1C , the gradient can also be indexed based on whether the gradient corresponds to the audio data 101 and / or the image 102.

[0033] Returning to Figure 1A, the client device 110 can then transmit the gradient 104 to the remote system 160 via one or more wired or wireless networks (e.g., the Internet, WAN, LAN, PAN, Bluetooth, and / or other networks). In some embodiments, the client device 110 can transmit the gradient 104 to the remote system 160 in response to determining that one or more conditions are met. The one or more conditions can include, for example, the client device 110 has authorized the transmission of the gradient 104, the client device 110 is charging, the client device 110 has at least a threshold charge state, the temperature of the client device 110 (based on one or more temperature sensors on the device) is below a threshold, the client device is not being held by a user, a time condition associated with the client device (e.g., between specific time periods, every N hours, etc. (where N is a positive integer) and / or other time conditions associated with the client device), whether a given number of thresholds of gradients have been generated by the client device, and / or other conditions. In this way, while the client device 110 monitors the satisfaction of one or more conditions, the client device 110 can generate multiple gradients. In some versions of those embodiments, in response to the remote system 160 requesting the gradient 104 from the client device 110 and in response to the client device 110 determining that one or more conditions are met, the client device 110 can transmit the gradient 104 to the remote system 160. In other embodiments, the client device 110 can transmit the gradient 104 to the remote system 160 in response to generating the gradient 104.

[0034] In some additional and / or alternative embodiments, the unsupervised learning engine 126 can also optionally provide the generated gradient 104 to the local training engine 132A. The local training engine 132A uses the gradient 104 to update one or more local machine learning models or portions thereof stored in the local machine learning model database 152A when it receives the generated gradient 104. For example, the local training engine 132A can update one or more weights of one or more local machine learning models (or subsets of their layers) used in generating the gradient 104, as described in more detail herein (e.g., with respect to Figure 1B and 1C ). Note that in some embodiments, the local training engine 132A can utilize batch techniques to update one or more local machine learning models or portions thereof based on the gradient 104 and additional gradients determined locally at the client device 110 based on other audio data and / or other image data.

[0035] As described above, the client device 110 may transmit the generated gradient 104 to the remote system 160. When the remote system 160 receives the gradient 104, the remote training engine 162 of the remote system 160 uses the gradient 104 and additional gradients 105 from multiple additional client devices 170 to update one or more weights of the global machine learning model layer stored in the global machine learning model database 152B. Each of the additional gradients 105 from the multiple additional client devices 170 may be generated based on the same or similar techniques as described above with respect to generating the gradient 104 (e.g., with respect to Figure 1B and 1C described), but based on audio data and / or image data locally generated at a corresponding one of the multiple additional client devices.

[0036] In some embodiments, the remote system 160 may also generate remote gradients based on publicly available data stored in one or more publicly available data databases 180. The one or more publicly available data databases 180 may include any repository of data that is publicly available via one or more networks and includes audio data, video data, and / or image data. For example, the one or more publicly available data databases 180 may be an online video sharing platform, an image sharing platform, and / or an audio sharing platform, which are not subject to access restrictions like the audio data 101 and image data 102 locally generated and / or locally stored at the client device 110. Each remote gradient may be generated remotely at the remote system 160 based on the same or similar techniques as described above with respect to generating the gradient 104 (e.g., with reference to Figure 1B and 1C described), but based on publicly available data retrieved from the one or more publicly available data databases 180. The remote gradients may also be used to update one or more weights of the global machine learning model layer stored in the global machine learning model database 152B.

[0037] In some versions of those embodiments, the remote training engine 162 can utilize the gradients 104, additional gradients 105, and remote gradients to update one or more weights of the global machine learning model layer. The remote system 160 can allocate the gradients 104, 105, and / or remote gradients to a particular iteration of updating the global machine learning model layer based on one or more criteria. The one or more criteria can include, for example, a threshold number of the gradients 104, 105, and / or remote gradients, a threshold duration of using the gradients and / or additional gradients for updating, and / or other criteria. Specifically, the remote training engine 162 can identify multiple subsets of gradients generated by the client device 110 and the multiple additional client devices 170 based on access-restricted data (e.g., audio data 101 and / or image data 102), and can identify multiple subsets of gradients generated by the remote system 160 based on publicly available data. Additionally, the remote training engine 162 can iteratively update the global machine learning model layer based on these subsets of gradients. For example, assume that the remote training engine 162 identifies a first subset and a third subset of gradients that only include the gradients 104, 105 locally generated at the client device 110 and the multiple additional client devices, and further assume that the remote training engine 162 identifies a second subset and a fourth subset of gradients that only include the remote gradients remotely generated at the remote system 160. In this example, the remote training engine 162 can update one or more weights of the global machine learning model layer based on the first subset of gradients locally generated at the client device, then update one or more weights of the global machine learning model layer based on the second subset of gradients remotely generated at the remote system 130, then update based on the third subset, then update based on the fourth subset, and so on.

[0038] In some further versions of those embodiments, the number of gradients included only in the subset of gradients locally generated by the client device 110 and the plurality of additional client devices may be the same as the additional number of remote gradients included only in the subset of gradients remotely generated at the remote system 160. In other versions of these embodiments, the remote system 160 may allocate the gradients 104 and additional gradients 105 to various subsets for a particular iteration based on the client devices 110, 170 transmitting the gradients. Each subset may optionally include gradients from at least one unique client device not included in another subset. For example, the remote training engine 162 may update one or more weights of a global machine learning model layer based on 50 gradients (or a first subset thereof) locally generated at the client device, then update one or more weights of the global machine learning model layer based on 50 gradients remotely generated at the remote system 160, then update one or more weights of the global machine learning model layer based on 50 gradients (or a second subset thereof) locally generated at the client device, and so on. In other further versions of those embodiments, the number of gradients included only in the subset of gradients locally generated by the client device 110 and the plurality of additional client devices may be different from the additional number of remote gradients included only in the subset of gradients remotely generated at the remote system 160. For example, the remote training engine 162 may update one or more weights of a global machine learning model layer based on 100 gradients locally generated at the client device, then update one or more weights of the global machine learning model layer based on 50 gradients remotely generated at the remote system 160, and so on. In other versions of these embodiments, the number of gradients included only in the subset of gradients locally generated by the client device 110 and the plurality of additional client devices and the additional number of remote gradients included only in the subset of gradients generated at the remote system 160 may vary throughout the training process. For example, the remote training engine 162 may update one or more weights of a global machine learning model layer based on 100 gradients locally generated at the client device, then update one or more weights of the global machine learning model layer based on 50 gradients remotely generated at the remote system 160, then update one or more weights of the global machine learning model layer based on 75 gradients locally generated at the client device, and so on.

[0039] In some embodiments, after updating one or more weights of the global machine learning model layer at the remote system 160, the supervised learning engine 164 may combine the updated global machine learning model layer with additional layers also stored in the global machine learning model database 152B to generate the combined machine learning model 106. For example, as Figure 1DAs shown, the remote training engine 162 can generate an updated global machine learning model layer 152B1 using the gradient 104, additional gradient 105, and remote gradient 203. Notably, the gradient 104 and additional gradient 105 can be stored in the first buffer 204 and the remote gradient 203 can be stored in the second buffer 205. The gradients can optionally be stored in the first buffer 204 and the second buffer 205 until a sufficient number of gradients appear in the remote system to identify multiple subsets of the gradients, as described above (e.g., with respect to the remote training engine 162). By updating the global machine learning model layer using the remote gradient 205 generated based on publicly available data and the gradients 104 and 105 generated based on the user's spoken utterances, different data is used to update the global machine learning model layer, resulting in a more robust global machine learning model layer than updating the global machine learning model layer based only on the user's spoken utterances. For example, the global machine learning model layer can be updated to generate a richer speech feature representation because it is not biased towards the spoken utterances typically received at the client devices 110, 170.

[0040] The supervised learning engine 164 can combine the updated global machine learning model layer 152B1 with an additional layer 152B2 that is an upstream layer of the updated global machine learning model layer 152B1, and train the updated global machine learning model layer 152B1 together with the additional layer 152B2, thereby generating a combined machine learning model 106. More specifically, the supervised learning engine 164 can connect at least one output layer of the updated global machine learning model layer 152B1 to at least one input layer of the additional layer 152B2. When connecting at least one output layer of the updated global machine learning model layer 152B1 to at least one input layer of the additional layer 152B2, the supervised learning engine 164 can ensure that the size and / or dimension of the output layer of the updated global machine learning model layer 152B1 is compatible with at least one input layer of the additional layer 152B2.

[0041] In addition, the supervised learning engine 164 can identify a plurality of training instances stored in the training instance database 252. Each of the plurality of training instances can include a training instance input and a corresponding training instance output (e.g., a ground truth output). The plurality of training instances identified by the supervised learning engine 164 can be based on the ultimate goal of training a given combined machine learning model. As a non-limiting example, assume that a given combined machine learning model is trained as an automatic speech recognition model. In this example, the training instance input of each training instance can include training audio data, and the training instance output can be the ground truth output corresponding to the phoneme or predicted token (corresponding to the training audio data). As another non-limiting example, assume that a given combined machine learning model is trained as an object classification model. In this example, the training instance input of each training instance can include training image data, and the training instance output can be the ground truth output corresponding to the object classification of the object included in the image data).

[0042] In addition, the training instance engine 164A of the supervised learning engine 164 can train the updated global machine learning model layer 152B1 and the additional layer 152B2 in a supervised manner based on the training instances 252. The error engine 164B of the supervised learning engine 164 can determine an error based on the training, and the backpropagation engine 164A can backpropagate the determined error across the additional layer 152B2 and / or update the weights of the additional layer 152B2, thereby training the combined machine learning model 106. In some embodiments, the updated global machine learning model layer 152B1 remains fixed as the determined error is backpropagated across the additional layer 152B2 and / or the weights of the additional layer 152B2 are updated. In other words, only the additional layer 152B2 is trained using supervised learning, while the updated global machine learning model layer 152B1 is updated using unsupervised learning. In some additional and / or alternative embodiments, the determined error is backpropagated across the updated global machine learning model layer 152B1 and the additional layer 152B2. In other words, unsupervised learning can be used to update the global machine learning model layer 152B1, and supervised learning can be used to train the combined machine learning model 105 including the updated global machine learning model layer 152B1.

[0043] Back to Figure 1A, the update distribution engine 166 can transmit the combined machine learning model 106 to the client device 110 and / or one or more of the plurality of additional client devices 170 in response to one or more of the client device 110 or the plurality of additional client devices 170 meeting one or more conditions. The one or more conditions can include, for example, a threshold duration and / or amount of training since the last update of weights and / or the update of the speech recognition model. The one or more conditions can additionally or alternatively include, for example, a measured improvement in the combined machine learning model 106 and / or a threshold duration has elapsed since the combined machine learning model 106 was last transmitted to the client device 110 and / or one or more of the plurality of additional client devices 170. When the combined machine learning model 106 is transmitted to the client device 110 (and / or one or more of the plurality of additional client devices 170), the client device 110 can store the combined machine learning model 106 in the local machine learning model database 152A, thereby replacing the previous version of the combined machine learning model 106. The client device 110 can then use the combined machine learning model 105 to make predictions based on further user input detected at the client device 110 (e.g., as described in more detail with respect to Figure 2 ). In some embodiments, the remote system 160 can also cause the updated global machine learning model layer to be transmitted to the client device 110 and / or the additional client device 170 together with the combined machine learning model 106. The client device 110 can store the updated global machine learning model layer in the local machine learning model database 152A, and can utilize the updated global machine learning model layer as part of the encoding for further audio data and / or further image data generated at the client device 110 after the audio data 101 and / or the image data 102 (e.g., used by the encoding engine 122A of Figure 1B and 1C ).

[0044] The client device 110 and the plurality of additional client devices 170 can continue to generate further gradients in the manner described herein and transmit the further gradients to the remote system 160. Additionally, the remote system 160 can continue to update the global machine learning model layer as described herein. In some embodiments, the remote system 160 can swap out the updated global machine learning model layer 152B1 in a given one of the combined machine learning models 106 with a further updated global machine learning model layer and can transmit the updated combined machine learning model to the client device. In some versions of those embodiments, the updated combined machine learning model can be further trained in a supervised manner as described herein before being transmitted to the client device. Thus, the combined machine learning models stored in the local machine learning model database 152A can reflect those generated and trained at the remote system 160, as described herein.

[0045] In some additional and / or alternative embodiments, the client device 110 and / or the additional client devices 170 can be limited to those of various institutions across different service sectors (e.g., medical institutions, financial institutions, etc.). In some versions of those embodiments, the combined machine learning model 106 can be generated in a manner that utilizes the underlying data across these different service sectors while also preserving the privacy of the underlying data used to generate the gradients 104 and / or additional gradients 105 for updating the global machine learning model layer. Additionally, the combined machine learning models 106 generated for these institutions can be stored remotely at the remote system 160 and accessed at the remote system 160 as needed.

[0046] By first training the global machine learning model layer in this manner and then combining it with additional layers of other machine learning models, the resulting combined machine learning model can be trained more effectively. For example, it can be trained to achieve the same performance level with fewer training instances. For example, a high level of accuracy can be achieved for broad coverage of dictated utterances. Additionally, by training the global machine learning model layer in this manner using unsupervised learning, the training of the global machine learning model layer can be computationally more efficient in the sense that it does not require labeled training instances.

[0047] Turning now Figure 2 to, the client device 110 is shown in the following embodiment where various on-device machine learning engines using the combined machine learning model described herein are included as part of (or communicate with) the automated assistant client 240. The corresponding machine learning models are also illustrated as docking with the various on-device machine learning engines. For simplicity, other components of the client device 210 are not shown in Figure 2 . Figure 2Illustrated is an example of how the automated assistant client 240 utilizes machine learning engines on various devices and their corresponding combined machine learning models when performing various actions.

[0048] Figure 2 The client device 210 in [description] is illustrated as having one or more microphones 211, one or more speakers 212, one or more visual components 213, and a display 214 (e.g., a touch-sensitive display). The client device 210 may also include pressure sensors, proximity sensors, accelerometers, magnetometers, and / or other sensors for generating other sensor data that supplements the audio data captured by one or more microphones 211. The client device 210 at least selectively executes the automated assistant client 240. In Figure 2 the example of [description], the automated assistant client 240 includes: a hotword detection engine 222, a non-hotword invocation engine 224, a continue conversation engine 226, a speech recognition engine 228, an object detection engine 230, an object classification engine 232, a natural language understanding (“NLU”) engine 234, and a fulfillment engine 236. The automated assistant client 240 also includes a voice capture engine 216 and a visual capture engine 218. The automated assistant client 240 may also include additional and / or alternative engines, such as a voice activity detector (VAD) engine, an endpoint detector engine, a lip movement engine, and / or other engines, as well as the associated machine learning models.

[0049] One or more cloud-based automated assistant components 280 may optionally be implemented on one or more computing systems (collectively referred to as “cloud” computing systems) that are communicatively coupled to the client device 210 via one or more local area networks and / or wide area networks (e.g., the Internet) generally indicated at 299. The cloud-based automated assistant components 280 may be implemented, for example, by a high-performance server cluster. In various embodiments, an instance of the automated assistant client 240, through its interaction with one or more cloud-based automated assistant components 280, may form a logical instance of an automated assistant 295 that appears to the user to be something with which the user can perform human-computer interactions (e.g., verbal interactions, gesture-based interactions, and / or touch-based interactions).

[0050] The client device 210 can be, for example: a desktop computing device, a laptop computing device, a tablet computing device, a mobile phone computing device, a computing device of a user's vehicle (e.g., an in-vehicle communication system, an in-vehicle entertainment system, an in-vehicle navigation system), a stand-alone interactive speaker, a smart appliance (such as a smart TV (or a standard TV equipped with a network dongle having an automated assistant function)) and / or a wearable device of the user including a computing device (e.g., a user's watch having a computing device, a user's glasses having a computing device, a virtual or augmented reality computing device). Additional and / or alternative client devices can be provided.

[0051] One or more vision components 213 can take various forms, such as a monocular camera, a stereo camera, a LIDAR component (or other laser-based component), a radar component, etc. For example, the vision capture engine 216 can use one or more vision components 213 to capture image data corresponding to a vision frame (e.g., an image frame, a laser-based vision frame) of the environment in which the client device 210 is deployed. In some embodiments, such vision frames can be used to determine whether a user is present near the client device 210 and / or the distance of the user (e.g., the user's face) relative to the client device 210. Such determination can be used, for example, to determine whether to activate Figure 2 the machine learning engine and / or other engines on the various devices described in. Additionally, the voice capture engine 218 can be configured to capture the spoken words of the user and / or other audio data captured via one or more microphones 211. Further, the client device 210 can include pressure sensors, proximity sensors, accelerometers, magnetometers, and / or other sensors for generating other sensor data that supplements the audio data captured via the microphone 211.

[0052] As described herein, such audio data and other non-microphone sensor data can be processed by Figure 2 the various engines depicted in for use at the client device 210 with reference to Figure 1A-1DPredictions are made using the corresponding combined machine learning models (including the updated global machine learning model layer) generated by the above-described manner. As some non-limiting examples, the hotword detection engine 222 may utilize the combined hotword detection model 222A to predict whether audio data includes a hotword for invoking the automated assistant 295 (e.g., "Ok Google", "Hey Google", "What is the weather Google?", etc.); the no-hotword invocation engine 224 may utilize the combined no-hotword invocation model 224A to predict whether non-microphone sensor data (e.g., image data) includes a gesture or signal for invoking the automated assistant 295 (e.g., based on the user's gaze and optionally further based on the user's mouth movement); the continue conversation engine 226 may utilize the combined continue conversation model 226A to predict whether further audio data includes an utterance directed to the automated assistant 295 (e.g., or to an additional user in the environment of the client device 210); the speech recognition engine 228 may utilize the combined speech recognition model 228A to predict phonemes and / or tokens corresponding to the audio data detected at the client device 210; the object detection engine 230 may utilize the combined object detection model 230A to predict the object locations in the image data included in the images captured at the client device 210; and the object classification engine 232 may utilize the combined object classification model 232A to predict the object classification of the objects in the image data included in the images captured at the client device 210.

[0053] In some embodiments, the client device 210 may further include an NLU engine 234 and an execution engine 236. The NLU engine 234 may perform on-device natural language understanding on the predicted phonemes and / or tokens generated by the speech recognition engine 228 using an NLU model 234A to generate NLU data. For example, the NLU data may include the intent corresponding to the spoken utterance and optional intent parameters (e.g., slot values). Additionally, the execution engine may utilize an on-device execution model 146A and generate execution data based on the NLU data. The execution data may define a local and / or remote response (e.g., answer) to the spoken utterance, an interaction with a locally installed application performed based on the spoken utterance, a command transmitted (directly or through a corresponding remote system) to an Internet of Things (IoT) device based on the spoken utterance, and / or other resolution actions performed based on the spoken utterance. The execution data is then provided for local and / or remote implementation / execution of a determined action to resolve the spoken utterance. The execution may include, for example, rendering a local and / or remote response (e.g., visual and / or auditory rendering (optionally using a local text-to-speech module)), interacting with a locally installed application, transmitting a command to an IoT device, and / or other actions. In other embodiments, the NLU engine 234 and the execution engine 236 may be omitted, and the speech recognition engine 228 may directly generate execution data based on the audio data. For example, assume that the speech recognition engine 228 processes the spoken utterance "turn on the lights" using a combined speech recognition model 228A. In this example, the speech recognition engine 228 may generate a semantic output and then transmit it to a software application associated with the lights, indicating that the lights should be turned on.

[0054] Notably, the cloud-based automated assistant component 280 includes cloud-based counterparts of the engines and models described herein. However, in various embodiments, these engines and models may not be invoked because the engines and models may be directly transferred to the client device 210 and executed locally on the client device 210, as described above with respect to Figure 2 In some embodiments, the client device 210 may further include an NLU engine 234 and an execution engine 236. The NLU engine 234 may perform on-device natural language understanding on the predicted phonemes and / or tokens generated by the speech recognition engine 228 using an NLU model 234A to generate NLU data. For example, the NLU data may include the intent corresponding to the spoken utterance and optional intent parameters (e.g., slot values). Additionally, the execution engine may utilize an on-device execution model 146A and generate execution data based on the NLU data. The execution data may define a local and / or remote response (e.g., answer) to the spoken utterance, an interaction with a locally installed application performed based on the spoken utterance, a command transmitted (directly or through a corresponding remote system) to an Internet of Things (IoT) device based on the spoken utterance, and / or other resolution actions performed based on the spoken utterance. The execution data is then provided for local and / or remote implementation / execution of a determined action to resolve the spoken utterance. The execution may include, for example, rendering a local and / or remote response (e.g., visual and / or auditory rendering (optionally using a local text-to-speech module)), interacting with a locally installed application, transmitting a command to an IoT device, and / or other actions. In other embodiments, the NLU engine 234 and the execution engine 236 may be omitted, and the speech recognition engine 228 may directly generate execution data based on the audio data. For example, assume that the speech recognition engine 228 processes the spoken utterance "turn on the lights" using a combined speech recognition model 228A. In this example, the speech recognition engine 228 may generate a semantic output and then transmit it to a software application associated with the lights, indicating that the lights should be turned on. Figure 1A-1DAs described. Nevertheless, a remote execution module may optionally be included that performs remote execution based on locally or remotely generated NLU data and / or fulfillment data. Additional and / or alternative remote engines may be included. As described herein, in various embodiments, device - based speech processing, device - based image processing, device - based NLU, device - based fulfillment, and / or device - based execution may be preferred at least because of the latency and / or reduced network usage they provide (due to not requiring client - server round - trips to parse spoken utterances). However, one or more cloud - based automated assistant components 280 may be used at least selectively. For example, such components may be used in parallel with device - based components and the output from such components may be used when local components fail. For example, if any of the device - based engines and / or models fail (e.g., due to the relatively limited resources of client device 110), the more robust resources of the cloud may be utilized.

[0055] Figure 3 A flowchart of an example method 300 is described, the method: generating gradients locally at a client device using unsupervised learning; transmitting the generated gradients to a remote system that utilizes the generated gradients to update the weights of a global machine - learning model layer; and, receiving at the client device a combined machine - learning model that includes the updated global machine - learning model layer and additional layers. For convenience, the operations of method 300 are described with reference to the system performing the operations. The system of method 300 includes one or more processors and / or other components of the client device. Additionally, although the operations of method 300 are shown in a particular order, this is not meant to be limiting. One or more operations may be reordered, omitted, or added.

[0056] In block 352, the system detects at the client device sensor data that captures one or more environmental attributes of the client device's environment. In some embodiments, block 352 includes an optional sub - block 352A. In the optional sub - block 352A, the system detects audio data that captures spoken utterances in the client device's environment via one or more microphones of the client device. The audio data captures at least a portion of the spoken utterances of the user of the client device. In other embodiments, block 352 includes an optional sub - block 352B. In the optional sub - block 352B, the system detects non - microphone sensor data via non - microphone sensors of the client device. The non - microphone sensor data may include, for example, image data that captures the client device's environment via the visual components of the client device.

[0057] At block 354, the system processes the sensor data using a local machine learning model stored locally at the client device to generate a prediction output. The local machine learning model at least includes a portion for generating an encoding of the sensor data. In some embodiments, the local machine learning model may also include an additional portion for decoding the encoding of the sensor data (e.g., as described in more detail above with respect to Figure 1B ). In some additional and / or alternative embodiments, the local machine learning model may also include an additional portion for making predictions based on the encoding of the sensor data (e.g., as described in more detail above with respect to Figure 1C ).

[0058] At block 356, the system generates a gradient based on the prediction output using unsupervised learning local to the client device. For example, assume that the sensor data detected at the client device is image data captured by a visual component of the client device, and an additional portion of the local machine learning model seeks to reconstruct the image data based on the encoding of the image data, thereby producing predicted image data. In this example, the system may compare the image data with the predicted audio data to determine the difference between them, and the system may generate a gradient based on the determined difference. As another example, assume that the sensor data detected at the client device is audio data captured by a microphone of the client device, including a first portion and a second portion after the first portion, and an additional portion of the local machine learning model seeks to predict the second portion of the audio data based on the encoding of the first portion of the audio data, thereby producing a predicted second portion of the audio data. In this example, the system may compare the second portion of the audio data with the predicted second portion of the audio data to determine the difference between them, and the system may generate a gradient based on the determined difference. Generating gradients using unsupervised learning at the client device is described in more detail herein (e.g., with respect to Figure 1B and 1C )

[0059] At block 358, the system determines whether conditions for transmitting the gradients generated at block 356 are met. The conditions can include, for example, that the client device is charging, the client device has at least a threshold charge state, the temperature of the client device (based on temperature sensors on one or more devices) is less than a threshold, the client device is not held by a user, a time condition associated with the client device (e.g., between specific time periods, every N hours, etc., where N is a positive integer) and / or other time conditions associated with the client device, whether a given one of the client devices has generated a threshold number of gradients, and / or other conditions. If, in an iteration of block 358, the system determines that the conditions for transmitting the gradients generated at block 356 are not met, the system can continue to monitor at block 358 whether the conditions are met. It is noted that when the system monitors the satisfaction of the conditions at block 358, the system can continue to generate additional gradients according to blocks 352 - 356 of method 300. If, in an iteration of block 358, the system determines that the conditions for transmitting the gradients generated at block 356 are met, the system can proceed to block 360.

[0060] At block 360, the system transmits the generated gradients from the client device to the remote system so that the remote system can use the generated gradients to update the weights of the globally stored machine learning model layer in the remote system. Additionally, multiple additional client devices can generate additional gradients according to method 300 and can transmit the additional gradients to the remote system when corresponding conditions are met at the additional client devices. Updating the weights of the globally stored machine learning model layer is described in more detail herein (e.g., with respect to Figure 1A and 4 ).

[0061] At block 362, the system receives, at the client device, a combined machine learning model from the remote system, which combined machine learning model includes the updated globally stored machine learning model layer and additional layers. It is noted that there is no arrow connecting blocks 360 and 362. This indicates that when the remote system determines, based on one or more conditions being met at the client device and / or the remote system, to transmit the combined machine learning model to the client as described in more detail herein (e.g., with respect to Figure 1A 、 1D and 4), the combined machine learning model is received at the client device.

[0062] At block 364, the system makes at least one prediction using the combined machine learning model based on user input detected at the user's client device. The prediction made at the client device may depend on the additional layers used to generate the combined machine learning model and / or the training instances used to train the combined machine learning model at the remote system. Making predictions at the client device using the combined machine learning model is described in more detail herein (e.g., with respect to Figure 2 for various engines and models).

[0063] Figure 4 depicts a flowchart of an illustrative example method 400 that: updates weights of a global machine learning model layer based on gradients received from multiple client devices and / or generated at a remote system based on publicly available data; generates a combined machine learning model including the updated global machine learning model layer and additional layers; and transmits the combined machine learning model to one or more of the multiple client devices. For convenience, the operations of method 400 are described with reference to a system that performs the operations. The system of method 400 includes one or more processors and / or other components of the remote system. Additionally, although the operations of method 400 are shown in a particular order, this is not meant to be limiting. One or more operations may be reordered, omitted, or added.

[0064] At block 452, the system receives gradients at the remote system from multiple client devices, where the gradients are locally generated at a corresponding one of the multiple client devices based on unsupervised learning at the corresponding one of the multiple client devices. Specifically, audio data and / or image data locally generated at the corresponding one of the client devices can be processed using a corresponding local machine learning model locally stored on the corresponding one of the multiple client devices. The corresponding local machine learning model can include corresponding portions for encoding the audio data and / or image data generated. Generation of gradients in an unsupervised manner at a corresponding one of the client devices is described in more detail herein (e.g., with respect to Figure 1A-1C and 3).

[0065] At block 454, the system identifies multiple additional gradients generated based on publicly available data. The publicly available data can be retrieved from an online video sharing platform, an image sharing platform, and / or an audio sharing platform, which are not subject to access restrictions like the audio data and image data locally generated at a corresponding one of the client devices. These additional gradients are also referred to herein as remote gradients because they are generated at the remote system using publicly available data. In some embodiments, block 454 can include optional sub-blocks 454A, 454B, and 454C. In sub-block 454A, the system retrieves publicly available data from a database (e.g., Figure 1A the publicly available data database 180). In sub-block 454B, the system processes the publicly available data using the global machine learning model to generate a predicted output. Similar to the corresponding local machine learning models at each of the multiple client devices, the global machine learning model can also include corresponding portions for encoding the audio data and / or image data generated. In sub-block 454C, the system uses unsupervised learning to generate multiple additional gradients based on the predicted output generated in sub-block 454B. Thus, the additional gradients can be generated in a manner similar to the gradients received from the multiple client devices (e.g., as described herein with respect toFigure 1A-1D generate additional gradients in a similar manner (more detailed description).

[0066] At block 456, the system updates the weights of the globally machine - learned model layer stored remotely in the remote system based on the received gradients and / or additional gradients. The globally machine - learned model layer corresponds to the portion of the corresponding local machine - learned model used to generate the encoded audio data and / or image data. Notably, as indicated by the dashed line from block 456 to block 452, the system can repeat the operations of blocks 452, 454, and 456 until the update of the globally machine - learned model layer is complete. The system can determine that the update of the globally machine - learned model layer is complete based on, for example, the following: a threshold duration and / or a threshold number of gradients since the updated weights and / or the globally machine - learned model layer were last trained as part of the combined machine - learned model; a measured improvement to the globally machine - learned model layer; and / or a threshold duration has elapsed since the updated weights and / or the globally machine - learned model layer were last trained as part of the combined machine - learned model. Once the globally machine - learned model layer is updated, the system can then proceed to block 458.

[0067] At block 458, the system generates a combined machine - learned model at the remote system that includes the updated globally machine - learned model layer and an additional layer. The system can connect the output layer of the globally machine - learned model layer to at least one input layer of the additional layer and can ensure that the size and / or dimensions of the output layer of the globally machine - learned model layer match those of at least one input layer of the additional layer. At block 460, the system remotely trains the combined machine - learned model at the remote system using supervised learning. The system can identify training instances for a particular purpose (e.g., speech recognition, hot - word detection, object detection, etc.) and can train the combined machine - learned model based on the training instances. This is described in more detail herein (e.g., with respect to Figure 1A and 1D ) the generation of the combined machine - learned model and its training in a supervised manner.

[0068] At block 462, the system determines whether a condition for transmitting the combined machine - learned model trained at block 460 is met. The condition can be based on whether the client device is ready to receive the combined machine - learned model (e.g., as described above with respect to Figure 3The conditions described by block 358 (the same as those described by block 358), other conditions specific to the remote system (e.g., the performance of the combined machine learning model meets a performance threshold based on the performance of the combined machine learning model, training the combined machine learning model based on a threshold number of training instances, etc.), and / or some combination of these conditions. If the system determines in an iteration of block 462 that the conditions for transmitting the combined machine learning model trained in block 460 are not met, the system can continue to monitor in block 462 whether the conditions are met. It is worth noting that when the system monitors the satisfaction of the conditions in block 462, the system can continue to update the global machine learning model layer and train the combined machine learning model according to blocks 452 - 460 of method 300. If in an iteration of block 462, the system determines that the conditions for transmitting the combined machine learning model trained in block 460 are met, the system can proceed to block 464.

[0069] At block 464, the system can transmit the combined machine learning model from the remote system to one or more of the multiple client devices. The system can transmit the combined machine learning model to each of the multiple client devices that transmit gradients to the remote system, additional client devices other than the client devices that transmit gradients to the remote system, or a subset of those client devices that transmit gradients to the remote system. The transmission of the combined machine learning model is described in more detail herein (e.g., with respect to Figure 1A the update distribution engine 166).

[0070] Figure 5 FIG. depicts a flowchart of an illustrative example method 500 that: generates gradients at a client device using unsupervised learning; updates the weights of the global machine learning model layer based on the gradients; trains a combined machine learning model including the updated global machine learning model layer and additional layers; and uses the combined machine learning model on the client device to make predictions based on user input detected on the client device. For convenience, the operations of method 500 are described with reference to the system performing the operations. The system of method 500 includes one or more processors and / or other components of the client device and / or the remote system. Additionally, although the operations of method 500 are shown in a particular order, this is not meant to be limiting. One or more operations can be reordered, omitted, or added.

[0071] At block 552, the system detects sensor data at the client device that captures one or more environmental attributes of the environment of the client device. The sensor data can be audio data generated by one or more microphones of the client device and / or non - microphone data (e.g., image data) generated by other sensors of the client device (e.g., as described with respect to Figure 3as described by block 352). At block 554, the system uses a local machine learning model stored locally at the client device to process the sensor data to generate a prediction output. The local machine learning model includes a portion for encoding the sensor data detected at the client device at block 552. For example, the portion for encoding the sensor data can be the encoder portion of an encoder-decoder model, the generator model of a CycleGAN model, and / or other portions of other models capable of generating a feature representation of the sensor data detected at the client device. At block 556, the system generates gradients based on the prediction output using unsupervised learning local to the client device. As described herein with respect to Figure 1B and 1C , generating the gradients can be based on the prediction output. At block 558, the system transmits the generated gradients from the client device to the remote system. When a condition is met at the client device (e.g., as described by block 358 with respect to Figure 3 ), the system can transmit the gradients (and optionally other gradients generated at the client device in addition to the gradients).

[0072] At block 560, the system receives the generated gradients at the remote system. More specifically, the system can receive the gradients (and optionally other gradients generated at the client device in addition to the gradients) and additional gradients locally generated at a plurality of additional client devices. At block 562, the system updates the weights of the globally stored machine learning model layer at the remote system based on the received gradients and / or additional gradients. The additional gradients can also include remote gradients generated at the remote system based on publicly available data, as described herein in more detail (e.g., with respect to Figure 1A , 1D and 4).

[0073] At block 564, the system remotely uses supervised learning at a remote system to train a combined machine learning model that includes an updated global machine learning model layer and an additional layer. The system trains the combined machine learning model using labeled training instances. The labeled training instances identified for training the combined machine learning can be based on the additional layer combined with the updated global machine learning model layer. For example, if the additional layer is an additional layer for a speech recognition model, the identified training instances can be specific to training the speech recognition model. Conversely, if the additional layer is an additional layer for an object detection model, the identified training instances may be specific to training the object detection model. At block 566, the system transfers the combined machine learning model from the remote system to the client device. In some implementations, the system can also transfer the updated global machine learning model layer to the client device. The client device can then use the updated global machine learning model layer to generate an encoding of additional sensor data generated at the client device, thereby updating the local machine learning model used to generate the predictive output. At block 568, the system makes at least one prediction using the combined machine learning model based on user input detected at the client device. Some non-limiting examples of predictions made using various combined machine learning models are described in more detail herein (e.g., with respect to Figure 2 ).

[0074] Figure 6 FIG. 6 is a block diagram of an example computing device 610 that can optionally be used to perform one or more aspects of the techniques described herein. In some implementations, one or more of the client device, the cloud-based automated assistant component, and / or other components can include one or more components of the example computing device 610.

[0075] The computing device 610 generally includes at least one processor 614 that communicates with a plurality of peripheral devices via a bus subsystem 612. These peripheral devices can include a storage subsystem 624 (including, for example, a memory subsystem 625 and a file storage subsystem 626), a user interface output device 620, a user interface input device 622, and a network interface subsystem 616. The input and output devices allow a user to interact with the computing device 610. The network interface subsystem 616 provides an interface to an external network and is coupled to corresponding interface devices in other computing devices.

[0076] The user interface input device 622 can include a keyboard, a pointing device such as a mouse, trackball, touchpad, or graphics tablet, a scanner, a touch screen incorporated into the display, an audio input device such as a voice recognition system, microphone, and / or other types of input devices. In general, the use of the term "input device" is intended to include all possible types of devices and ways of inputting information into the computing device 510 or onto a communication network.

[0077] The user interface output device 620 may include a display subsystem, a printer, a fax machine, or a non-visual display such as an audio output device. The display subsystem may include a cathode ray tube (CRT), a flat panel device such as a liquid crystal display (LCD), a projection device, or some other mechanism for generating a visible image. The display subsystem may also provide a non-visual display, for example, via an audio output device. In general, the use of the term "output device" is intended to include all possible types of devices and ways of outputting information from the computing device 610 to a user or another machine or computing device.

[0078] The storage subsystem 624 stores programming and data constructs that provide the functionality of some or all of the modules described herein. For example, the storage subsystem 624 may include logic for performing selected aspects of the methods disclosed herein and for implementing Figure 1A and 1B the various components depicted in

[0079] These software modules are typically executed by the processor 614, alone or in combination with other processors. The memory 625 used in the storage subsystem 624 may include multiple memories, including a main random access memory (RAM) 630 for storing instructions and data during program execution and a read-only memory (ROM) 632 in which fixed instructions are stored. The file storage subsystem 626 may provide permanent storage for program and data files and may include a hard disk drive, a floppy disk drive together with an associated removable medium, a CD-ROM drive, an optical drive, or a removable media cartridge. The modules implementing the functionality of certain embodiments may be stored in the storage subsystem 624 by the file storage subsystem 626, or in other machines accessible by the processor 614.

[0080] The bus subsystem 612 provides a mechanism for enabling the various components and subsystems of the computing device 610 to communicate with each other as expected. Although the bus subsystem 612 is schematically shown as a single bus, alternative embodiments of the bus subsystem may use multiple buses.

[0081] The computing device 610 can be of different types, including workstations, servers, computing clusters, blade servers, server farms, or any other data processing system or computing device. Due to the ever-changing nature of computers and networks, for the purposes of illustrating some embodiments, Figure 6 the description of the computing device 610 depicted in Figure 6 is only intended as a specific example. Many other configurations of the computing device 610 may have more or fewer components than the computing device depicted in

[0082] In the case where the systems and methods discussed herein collect or otherwise monitor personal information about a user or can utilize personal and / or monitored information), the following opportunities may be provided to the user: to control whether a program or feature collects user information (e.g., information about the user's social network, social actions or behaviors, occupation, user preferences, or the user's current location), or to control whether and / or how to receive content from a content server that may be more relevant to the user. Additionally, certain data may be processed in one or more ways before storage or use to remove personally identifiable information. For example, a user's identity may be processed such that personally identifiable information cannot be determined for the user, or the user's geographical location may be generalized (e.g., generalized to a city, zip code, or state level) in the case where location information is obtained, such that the user's specific location cannot be determined. Thus, the user can control how information about the user is collected and used.

[0083] In some embodiments, a method executed by one or more processors of a client device is provided and includes: detecting audio data via one or more microphones of the client device, the audio data capturing at least a portion of the spoken words of a user of the client device, and processing the audio data using a local machine learning model locally stored on the client device to generate a prediction output. The local machine learning model includes a portion for generating an encoding of the audio data. The method further includes generating a gradient based on the prediction output using unsupervised learning, and transmitting the generated gradient from the client device to a remote system such that the remote system utilizes the generated gradient to update weights of a global machine learning model layer, the weights being remotely stored in the remote system and corresponding to the portion of the local machine learning model for generating the encoding of the audio data. The method further includes, after the remote system updates the weights of the global machine learning model layer using the generated gradient and the remote further updates the weights based on additional gradients from additional client devices, receiving, at the client device, a combined machine learning model from the remote system, the combined machine learning model including the updated global machine learning model layer and one or more additional layers, and making at least one prediction using the combined machine learning model based on further audio data detected via one or more of the microphones of the client device and capturing at least a portion of further spoken words of the user of the client device.

[0084] These and other implementations of the technology may include one or more of the following features.

[0085] In some embodiments, the portion of the local machine learning model locally stored on the client device is used to generate the encoding. The encoding is processed using an additional portion of the local machine learning model to generate the prediction output, and the prediction output is predicted audio data. In some versions of those embodiments, generating the gradient based on the prediction output includes comparing the audio data that captures at least a portion of the spoken utterance of the user of the client device with the predicted audio data, and generating the gradient based on comparing the audio data and the predicted audio data. In other versions of those embodiments, the audio data that captures at least a portion of the spoken utterance of the user of the client device captures a first portion of the spoken utterance, followed by a second portion of the spoken utterance, and generating the gradient is based on comparing the predicted audio data with additional audio data corresponding to the second portion of the spoken utterance. In still other versions of those embodiments, comparing the predicted audio data with the additional audio data corresponding to the second portion of the spoken utterance includes comparing the audio waveform corresponding to the additional audio data with the predicted audio waveform corresponding to the predicted audio data, the additional audio data corresponding to the second portion of the spoken utterance, and generating the gradient based on comparing the audio waveform and the predicted audio waveform.

[0086] In some embodiments, the combined machine learning model is an automatic speech recognition ASR model, and the at least one prediction includes a plurality of predicted phonemes or a plurality of predicted tokens corresponding to the further spoken utterance.

[0087] In some embodiments, receiving the combined machine learning model and using the combined machine learning model is further after the remote system updates the weights of the global machine learning model layer using publicly available audio data to generate further gradients. Each of the further gradients is remotely generated on the remote server based on the following: unsupervised learning on the remote system, the unsupervised learning being based on processing the publicly available audio data that captures publicly available spoken utterances using the global machine learning model remotely stored on the remote server, and the global machine learning model including the global machine learning model layer for generating further encoding of the publicly available audio data. In some versions of those embodiments, the remote system updates the weights of the global machine learning model layer using the gradients generated at the client device and additional gradients generated using unsupervised learning at an additional client device, and after using the gradients generated at the client device and the additional gradients generated at the additional client device, the remote system updates the weights of the global machine learning model layer using the further gradients generated at the remote system.

[0088] In some embodiments, a method executed by one or more processors of a client device is provided and includes receiving gradients at a remote system from a plurality of client devices. Each of the gradients is locally generated at a corresponding one of the plurality of client devices based on: unsupervised learning at the corresponding one of the plurality of client devices, the unsupervised learning being based on processing audio data capturing an oral utterance using a respective local machine learning model locally stored on the corresponding one of the plurality of client devices, and the respective local machine learning model including a respective portion for generating an encoding of the audio data; the method further includes updating weights of a global machine learning model layer based on the received gradients, the weights being remotely stored at the remote system and corresponding to the portion of the encoding for generating the audio data of the respective local machine learning model. The method further includes, after updating the weights of the global machine learning model layer based on the generated gradients: generating a combined machine learning model including the updated global machine learning model layer and one or more additional layers, and training the combined machine learning model using supervised learning. The method further includes, after training the combined machine learning model: transmitting the combined machine learning model to one or more of the plurality of client devices. The one or more of the plurality of client devices make at least one prediction using the combined machine learning based on further audio data capturing at least a portion of a further oral utterance of a user of the client device detected via one or more of the microphones of the client device.

[0089] These and other implementations of the technology may include one or more of the following features.

[0090] In some embodiments, generating the combined machine learning model includes connecting an output layer of the updated global machine learning model layer to at least one input layer of the one or more additional layers.

[0091] In some embodiments, training the combined machine learning model using supervised learning includes identifying a plurality of training instances, each of the training instances having: a training instance input including training audio data, and a corresponding training instance output including a ground truth output. The method further includes determining an error based on applying the plurality of training instances as inputs across the combined machine learning model to generate corresponding predicted outputs and comparing the corresponding predicted outputs with the corresponding training instance outputs, and updating weights of the one or more additional layers of the combined machine learning model based on the error while keeping the updated global machine learning model layer of the combined machine learning model fixed.

[0092] In some embodiments, training the combined machine learning model using supervised learning includes identifying a plurality of training instances, each of the training instances having: a training instance input including training audio data, and a corresponding training instance output including a ground truth output. The method further includes determining an error based on applying the plurality of training instances as inputs across the combined machine learning model to generate corresponding predicted outputs and comparing the corresponding predicted outputs with the corresponding training instance outputs, and backpropagating the determined error across one or more of the additional layers of the combined machine learning model and one or more of the updated global machine learning model layers of the combined machine learning model.

[0093] In some embodiments, the method further includes receiving additional gradients at the remote system from the plurality of client devices. Each of the additional gradients is locally generated at the corresponding one of the plurality of client devices based on: unsupervised learning at the corresponding one of the plurality of client devices, the unsupervised learning being based on processing additional audio data that captures additional dictated utterances using a respective local machine learning model locally stored on the corresponding one of the plurality of client devices, and the respective local machine learning model includes the respective portion for generating the encoding of the additional audio data. The method further includes further updating weights of the global machine learning model layer remotely stored at the remote system based on the received additional gradients. The method further includes, after further updating the weights of the global machine learning model layer based on the received additional gradients: modifying the combined machine learning model to generate an updated combined machine learning model, the updated combined machine learning model including the further updated global machine learning model layer and one or more of the additional layers, and training the updated combined machine learning model using supervised learning. The method further includes, after training the updated combined machine learning model: transmitting the updated combined machine learning model to one or more of the plurality of client devices to replace the combined machine learning model.

[0094] In some embodiments, the method further includes retrieving publicly available audio data from one or more databases, the publicly available audio data capturing a plurality of publicly available dictated utterances, processing the publicly available audio data using a global machine learning model to generate a predicted output. The global machine learning model includes the global machine learning model layer for generating the corresponding encoding of the publicly available audio data. The method further includes using unsupervised learning to generate a plurality of additional gradients based on the predicted output, and further updating the weights of the global machine learning model layer remotely stored at the remote system is based on the plurality of additional gradients.

[0095] In some versions of those embodiments, updating the weights of the global machine learning model layer stored remotely at the remote system includes identifying a first set of gradients based on the gradients received from the plurality of client devices, identifying a second set of gradients based on the additional gradients generated based on the publicly available audio data, updating the global machine learning model layer based on the first set of gradients, and after updating the global machine learning model layer based on the first set of gradients: updating the global machine learning model layer based on the second set of gradients.

[0096] In yet other versions of those embodiments, identifying the first set of gradients is based on a first threshold number of the received gradients included in the first set of gradients, and identifying the second set of gradients is based on a second threshold number of the generated gradients included in the second set of gradients.

[0097] In some embodiments, a method performed by one or more processors of a client device is provided and includes: detecting, via one or more microphones of the client device, audio data capturing an oral utterance of a user of the client device, processing the audio data using a local machine learning model stored locally on the client device to generate a predicted output, wherein the local machine learning model includes a portion for generating an encoding of the audio data, locally generating, on the client device, a gradient based on the predicted output using unsupervised learning, transmitting the generated gradient from the client device to a remote system, receiving, at the remote system, the generated gradient from the client device, and updating, at the remote system, one or more weights of a global machine learning model layer based on the generated gradient. The method further includes, after updating one or more weights of the global machine learning model layer based on the generated gradient, remotely training, at the remote system, a combined machine learning model including the updated global machine learning model layer and one or more additional layers using supervised learning. The method further includes transmitting the combined machine learning model from the remote system to the client device, and making at least one prediction using the combined machine learning model based on further audio data of at least a portion of a further oral utterance of the user of the client device detected via one or more of the microphones of the client device.

[0098] These and other implementations of the technology may include one or more of the following features.

[0099] In some embodiments, the combined machine learning model is an automatic speech recognition (ASR) model, and the at least one prediction includes a plurality of predicted phonemes or a plurality of predicted tokens corresponding to the further oral utterance.

[0100] In some embodiments, the combined machine learning model is a hotword detection model, and the at least one prediction includes an indication of whether the further utterance includes a hotword that invokes an automated assistant.

[0101] In some embodiments, the combined machine learning model is a voice activity detection model, and the at least one prediction includes an indication of whether the further utterance is human speech in the environment of the client device.

[0102] Various embodiments may include a non-transitory computer-readable storage medium that stores instructions executable by one or more processors (e.g., a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), and / or a tensor processing unit (TPU)) to perform a method such as one or more of the methods described herein. Other embodiments may include an automated assistant client device (e.g., a client device including at least an automated assistant interface for docking with a cloud-based automated assistant component) that includes a processor operable to execute stored instructions to perform a method such as one or more of the methods described herein. Other embodiments may include a system of one or more servers that includes one or more processors operable to execute stored instructions to perform a method such as one or more of the methods described herein.

Claims

1. A method executed by one or more processors of a client device, the method comprising: Detect audio data via one or more microphones of the client device, the audio data capturing at least a portion of the spoken words of a user of the client device; Process the audio data using a local machine learning model locally stored on the client device to generate a prediction output, wherein the local machine learning model includes a first set of local machine learning model layers and a second set of local machine learning model layers, wherein the first set of local machine learning model layers is configured to generate an encoding of the audio data based on processing the audio data, and wherein the second set of local machine learning model layers is configured to generate the prediction output based on processing the encoding of the audio data generated using the first set of local machine learning model layers; Generate a gradient based on the prediction output using unsupervised learning; and Transmit the generated gradient from the client device to a remote system such that the remote system utilizes the generated gradient to update weights of a global machine learning model layer, the weights being remotely stored in the remote system and structurally corresponding to the first set of local machine learning model layers of the local machine learning model that generate the encoding of the audio data, and After the remote system updates the weights of the global machine learning model layer using the generated gradient received from the client device and additional gradients received from additional client devices: Receive, at the client device, a combined machine learning model from the remote system, the combined machine learning model including the updated global machine learning model layer and one or more additional layers; and Make at least one prediction using the combined machine learning model based on further audio data detected via one or more of the microphones of the client device that captures at least a portion of further spoken words of the user of the client device.

2. The method according to claim 1, wherein, The prediction output is predicted audio data.

3. The method according to claim 2, wherein, Generating the gradient based on the prediction output includes: Comparing the audio data capturing at least a portion of the spoken words of the user of the client device with the predicted audio data; and Generating the gradient based on comparing the audio data and the predicted audio data.

4. The method according to claim 2, wherein, The audio data capturing at least a portion of the spoken words of the user of the client device captures a first portion of the spoken words followed by a second portion of the spoken words, and wherein generating the gradient is based on comparing the predicted audio data with additional audio data corresponding to the second portion of the spoken words.

5. The method according to claim 4, wherein, Comparing the predicted audio data with the additional audio data corresponding to the second portion of the spoken words includes: Comparing an audio waveform corresponding to the additional audio data with a predicted audio waveform corresponding to the predicted audio data, the additional audio data corresponding to the second portion of the spoken words; and Generating the gradient based on comparing the audio waveform and the predicted audio waveform.

6. The method according to claim 1, wherein, The combined machine learning model is an automatic speech recognition (ASR) model, and wherein, the at least one prediction includes a plurality of predicted phonemes or a plurality of predicted tokens corresponding to the further dictated utterance.

7. The method according to any one of claims 1 to 6, wherein, After receiving the combined machine learning model and using the combined machine learning model to further update the weights of the global machine learning model layer using publicly available audio data in the remote system to generate further gradients, wherein each of the further gradients is remotely generated on a remote server based on: Unsupervised learning at the remote system, the unsupervised learning being based on processing the publicly available audio data capturing the publicly available dictated utterance using the global machine learning model remotely stored at the remote server, wherein the global machine learning model includes the global machine learning model layer for generating a further encoding of the publicly available audio data.

8. The method according to claim 7, wherein, The remote system updates the weights of the global machine learning model layer using the gradients generated at the client device and additional gradients generated using unsupervised learning at an additional client device, and wherein, after using the gradients generated at the client device and the additional gradients generated at the additional client device, the remote system updates the weights of the global machine learning model layer using the further gradients generated at the remote system.

9. A method executed by one or more processors of a remote system, the method comprising: Receiving gradients at the remote system from a plurality of client devices, wherein each of the gradients is locally generated at a corresponding one of the plurality of client devices based on: Unsupervised learning at the corresponding one of the plurality of client devices, the unsupervised learning being based on processing audio data capturing the dictated utterance of a user of the client device using a respective local machine learning model locally stored on the corresponding one of the plurality of client devices, wherein the respective local machine learning model includes a first set of local machine learning model layers and a second set of local machine learning model layers, wherein the first set of local machine learning model layers is for generating an encoding of the audio data based on processing the audio data, wherein the second set of local machine learning model layers is for generating a predicted output based on processing the encoding of the audio data generated using the respective first set of local machine learning model layers, and wherein the gradients are generated using unsupervised learning based on the predicted output; Updating the weights of the global machine learning model layer based on the received gradients, the weights being remotely stored in the remote system and structurally corresponding to the respective first set of local machine learning model layers of the respective local machine learning model for generating the encoding of the audio data; After updating the weights of the global machine learning model layer based on the generated gradients: Generating a combined machine learning model including the updated global machine learning model layer and one or more additional layers; and Training the combined machine learning model using supervised learning; and After training the combined machine learning model: Transmit the combined machine learning model to one or more of the plurality of client devices, wherein the one or more of the plurality of client devices make at least one prediction using the combined machine learning based on further audio data that is detected via one or more microphones of the client device and that captures at least a portion of further spoken utterances of the user of the client device.

10. The method according to claim 9, wherein, Generating the combined machine learning model includes connecting an output layer of the updated global machine learning model layer to at least one input layer of the one or more additional layers.

11. The method according to claim 9 or claim 10, wherein, Training the combined machine learning model using supervised learning includes: Identifying a plurality of training instances, each of the training instances having: A training instance input including training audio data, and A corresponding training instance output including a ground truth output; Determining an error based on applying the plurality of training instances as inputs across the combined machine learning model to generate corresponding predicted outputs and comparing the corresponding predicted outputs with the corresponding training instance outputs; and Updating weights of the one or more additional layers of the combined machine learning model based on the error while keeping the updated global machine learning model layer of the combined machine learning model fixed.

12. The method according to claim 9 or claim 10, wherein, Training the combined machine learning model using supervised learning includes: Identifying a plurality of training instances, each of the training instances having: A training instance input including training audio data, and A corresponding training instance output including a ground truth output; Determining an error based on applying the plurality of training instances as inputs across the combined machine learning model to generate corresponding predicted outputs and comparing the corresponding predicted outputs with the corresponding training instance outputs; and Backpropagating the determined error across one or more layers of the one or more additional layers of the combined machine learning model and one or more layers of the updated global machine learning model layer of the combined machine learning model.

13. The method according to claim 9, further comprising: Receiving additional gradients at the remote system from the plurality of client devices, wherein each of the additional gradients is locally generated at the corresponding one of the plurality of client devices based on: Unsupervised learning at the corresponding one of the plurality of client devices, the unsupervised learning being based on processing additional audio data that captures additional spoken utterances using a respective local machine learning model that is locally stored on the corresponding one of the plurality of client devices; Based on the received additional gradients, further updating weights of the global machine learning model layer remotely stored at the remote system; After further updating the weights of the global machine learning model layer based on the received additional gradients: Modifying the combined machine learning model to generate an updated combined machine learning model, the updated combined machine learning model including the further updated global machine learning model layer and one or more of the additional layers; and Train the updated combined machine learning model using supervised learning; and After training the updated combined machine learning model: Transmit the updated combined machine learning model to one or more of the multiple client devices to replace the combined machine learning model.

14. The method according to claim 9, further comprising: Retrieve publicly available audio data from one or more databases, the publicly available audio data capturing multiple publicly available spoken utterances; Process the publicly available audio data using a global machine learning model to generate a prediction output, wherein the global machine learning model includes the global machine learning model layer for generating the corresponding encoding of the publicly available audio data; Use unsupervised learning to generate multiple additional gradients based on the prediction output; and Wherein updating the weights of the global machine learning model layer remotely stored at the remote system is further based on the multiple additional gradients.

15. The method according to claim 14, wherein, Updating the weights of the global machine learning model layer remotely stored at the remote system includes: Identifying a first set of gradients according to the gradients received from the multiple client devices; Identifying a second set of gradients according to the additional gradients generated based on the publicly available audio data; Updating the global machine learning model layer based on the first set of gradients; and After updating the global machine learning model layer based on the first set of gradients: Updating the global machine learning model layer based on the second set of gradients.

16. According to the method of claim 15, wherein, Identifying the first set of gradients is based on a first threshold number of the received gradients included in the first set of gradients, and wherein identifying the second set of gradients is based on a second threshold number of the generated gradients included in the second set of gradients.

17. A method executed by one or more processors, the method comprising: Detect audio data capturing the spoken utterance of the user of the client device via one or more microphones of the client device; Process the audio data using a local machine learning model locally stored on the client device to generate a prediction output, Wherein the local machine learning model includes a first set of local machine learning model layers and a second set of local machine learning model layers, Wherein the first set of local machine learning model layers is used to generate an encoding of the audio data based on processing the audio data, and Wherein the second set of local machine learning model layers is used to generate the prediction output based on processing the encoding of the audio data generated using the first set of local machine learning model layers; Locally generate gradients on the client device using unsupervised learning based on the prediction output; Transmit the generated gradients from the client device to the remote system; Receive the generated gradients at the remote system from the client device; Update one or more weights of the global machine learning model layer at the remote system based on the generated gradients, wherein the global machine learning model layer structurally corresponds to the first set of local machine learning model layers of the local machine learning model; After updating one or more weights of the global machine learning model layer based on the generated gradients, remotely train, at the remote system, a combined machine learning model including the updated global machine learning model layer and one or more additional layers using supervised learning; Transmit the combined machine learning model from the remote system to the client device; and Use the combined machine learning model to make at least one prediction based on further audio data that captures at least a portion of further spoken utterances of the user of the client device detected via one or more of the microphones of the client device.

18. According to the method of claim 17, wherein, The combined machine learning model is an automatic speech recognition (ASR) model, and wherein the at least one prediction includes a plurality of predicted phonemes or a plurality of predicted tokens corresponding to the further spoken utterances.

19. According to the method of claim 17, wherein, The combined machine learning model is a hotword detection model, and wherein the at least one prediction includes an indication of whether the further spoken utterances include a hotword that invokes an automated assistant.

20. According to the method of claim 17, wherein, The combined machine learning model is a voice activity detection model, and wherein the at least one prediction includes an indication of whether the further spoken utterances are human speech in the environment of the client device.