Ephemeral and / or federated training of audio-based machine learning models from streams of audio data generated via radio stations

Ephemeral and federated learning techniques process audio data from radio stations to update ML models, addressing the limitations of traditional methods by enhancing model robustness and precision for diverse languages.

JP2025529844APending Publication Date: 2025-09-09GOOGLE LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2025510390
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-12-05
Filing Date
2022-12-06
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

Existing ML model updating techniques are limited by explicit user input and lack of diversity in data, particularly for tail languages, leading to slow updates and limited effectiveness in processing diverse languages.

Method used

Utilizing ephemeral and federated learning techniques to process audio data from radio stations worldwide, generating client and remote gradients to update a global ML model, especially for tail languages, with deduplication and unsupervised learning to enhance model robustness and precision.

Benefits of technology

Enables efficient updating of global ML models with diverse audio data, improving precision and recall for languages not well-defined by traditional methods, and preventing overfitting, thus making the model more robust and applicable to a broader user base.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025529844000001_ABST
    Figure 2025529844000001_ABST
Patent Text Reader

Abstract

Embodiments disclosed herein are directed to utilizing ephemeral and / or federated learning techniques to update audio-based machine learning (ML) model(s) based on processing streams of audio data generated via radio station(s) around the world. This enables the audio-based ML model(s) to learn and / or understand expressions for languages ​​around the world, including languages ​​with no or minimal audio data. In various embodiments, one or more deduplication techniques may be utilized to prevent the same stream of audio data from being over-utilized when updating the audio-based ML model(s). In various embodiments, a given client device may determine whether to employ ephemeral or federated learning techniques based, for example, on its connection status with a remote system. Generally, the stream of audio data is received at the client device, but ephemeral learning techniques may be implemented on the client device and / or on the remote system.
Need to check novelty before this filing date? Find Prior Art

Description

[Background technology]

[0001] Federated learning of machine learning (ML) model(s) is an increasingly popular ML technique for updating ML model(s). In traditional federated learning, an on-device ML model is stored locally on a user's client device, and a global ML model, which is a cloud-based counterpart of the on-device ML model, is stored remotely in a remote system (e.g., a cluster of servers). The client device can use the on-device ML model to process user data detected on the client device to generate predicted outputs and compare the predicted outputs with ground truth outputs to generate client gradients. The client device can then send the client gradients to a remote system. The remote system can update the weights of the global ML model using the client gradients, and optionally, additional client gradients generated in a similar manner on additional client devices. The remote system can then send the global ML model or updated weights of the global ML model to the client device. The client device can then replace the on-device ML model with the global ML model or replace the weights of the on-device ML model with the updated weights of the global ML model, thereby updating the on-device ML model.

[0002] Ephemeral learning of ML model(s) is another ML technique for updating ML model(s) that is becoming increasingly popular. In traditional ephemeral learning, an on-device ML model is also stored locally on the user's client device, and a global ML model is also stored remotely at a remote system. However, in contrast to traditional ephemeral learning, traditional ephemeral learning has a temporal component. For example, user data may be transmitted to a remote system to fulfill a particular fulfillment. While the user data is being processed by the remote system to fulfill a particular fulfillment, the remote system can also use a global ML model to process the user data received from the client device and generate predicted outputs, generate remote gradients based on at least the predicted outputs (e.g., using self-supervised or unsupervised learning techniques), and discard the user data without storing any of the user data in non-transitory storage at the remote system. This allows the remote system to update the global ML model based on the remote gradients while the received user data is temporarily available at the remote system in a privacy-conscious manner. Also, for example, user data may be processed locally on the client device to generate client gradients in the same or similar manner as described above for federated learning, but the client gradients may be transmitted to a remote system such that they are only available for updating the global ML model in an ephemeral manner.

[0003] However, scenarios for utilizing these different ML techniques to update ML model(s) are generally limited to updating the ML model(s) based on explicit user input provided by a user at each client device. As a result, updating ML model(s) using these ML techniques can take a long time. Furthermore, the user data utilized by these different ML techniques is generally limited to a small subset of user-provided voice utterances and / or commands (e.g., "Assistant, set my alarm for 6:00 AM") for which these ML model(s) are utilized, and the ML model(s) only target common or well-defined languages ​​(e.g., English, Spanish, etc.), but not tail languages. As a result, updating ML model(s) using these ML techniques may simply enhance the ML model(s) to process and / or understand well-known voice utterances and / or commands in well-known languages. Thus, there is a need in the art to extend the utility of using these different ML techniques beyond explicit user input and to broaden the diversity of data utilized by these different ML techniques. Summary of the Invention

[0004] Implementations described herein are directed to utilizing various privacy-aware machine learning (ML) techniques to update a global ML model based on processing audio data from radio stations around the world. In some implementations, a client device may receive a stream of audio data capturing a stream of speech utterances in a given language from a given radio station. The client device may process the stream of audio data using an on-device ML model, which is an on-device counterpart of the global ML model stored in the client device's on-device storage, and may generate client gradients based on processing the stream of audio data using unsupervised or self-supervised learning techniques and with the on-device ML model. In some variations of these implementations, in accordance with ephemeral learning ML techniques for updating the global ML model, the client device may synchronously transmit the client gradients to a remote system (e.g., a high-performance server or cluster of high-performance servers) and cause the remote system to update the global ML model for the given language based on the client gradients and, optionally, based on additional client gradients generated by additional client devices in the same or similar manner. In additional or alternative variations of these embodiments, in accordance with a transient learning ML technique for updating the global ML model, the client device may asynchronously send client gradients to the remote system and cause the remote system to update the global ML model for a given language based on the client gradients, and optionally based on additional client gradients generated in the same or similar manner by additional client devices. In additional or alternative embodiments, the client device (or other computing device) may send a stream of audio data directly to the remote system.The remote system may process the stream of audio data using the global ML model and may generate remote gradients based on processing the stream of audio data using unsupervised or semi-supervised learning techniques and using the global ML model. In these implementations, the remote system may update the global ML model for a given language based on the remote gradients and for the given language, and optionally based on other remote gradients that the remote system generated in the same or similar manner.

[0005] Thus, the techniques described herein may utilize a combination of ephemeral and federated learning techniques to update a global ML model based on streams of audio data from different radio stations that broadcast the streams in different languages, extending the utility of using these different ML techniques without being limited to explicit user input and broadening the diversity of data utilized by these different ML techniques.

[0006] For example, assume that a given user of a given client device is located in South Africa (a country located on the African continent with 11 official languages ​​and various dialects). Further assume that the only language that the global multilingual ASR ML model is trained to recognize among these 11 official South African languages ​​is English, which is not the most common language spoken in South Africa. As a result, this example global multilingual ASR model may not be useful to a large portion of the South African population. While a publicly available repository of audio and / or video data may contain streams of audio data capturing another 10 official languages ​​(and another non-official language), these streams of audio data for the other 10 languages ​​(and another non-official language) may be quite limited in volume and therefore insufficient to train the global multilingual ASR model. This is because most of the other 10 languages ​​(and other languages) are not widespread in other parts of the world and / or may be considered tail languages, spoken by only a small minority of people in the world, such as tribal languages ​​in South Africa that utilize click consonants. However, further assume that there are hundreds of South African radio stations that broadcast streams of audio data in all of these different South African languages. Thus, by using a combination of ephemeral and federated learning techniques described herein to update the global multilingual ASRML model based on the streams of audio data in all of the different languages ​​from these hundreds of radio stations, the global multilingual ASRML model can be trained to recognize all of these languages.

[0007] While the above example is described with respect to a global ML model that is a global multilingual ASR model, it should be understood that this is for purposes of illustration and not intended to be limiting. Rather, it should be understood that the techniques described herein can be utilized in updating any global ML model (also referred to herein as “audio-based ML model(s)”) that processes audio data, features generated based on processing the audio data, and / or generates synthesized audio data. For example, the techniques described herein may be utilized to update a global language representation ML model that extracts features from audio data to generate rich feature representations of speech utterances captured in the audio data, a global voice activity detection ML model that detects voice activity in the audio data, a global language identification ML model that detects a given language spoken in the audio data, a global natural language understanding (NLU) ML model that processes features generated based on processing the audio data to analyze the audio data, a global text-to-speech (TTS) ML model that generates synthesized speech audio data in a given language, and / or any other global ML model described herein.

[0008] In various implementations, a client device may determine whether to implement ephemeral learning techniques or federated learning techniques to generate client gradients for updating a global ML model for a given language. In some variations of these implementations, the client device may make this determination based on the connection status between the client device and the remote system. For example, in response to determining that the connection status indicates a strong or stable connection between the client device and the remote system (e.g., via a Wi-Fi network, a cellular network, etc.), the client device may implement the ephemeral learning technique, generate client gradients locally at the client device, and synchronously transmit the client gradients to the remote system. Additionally or alternatively, the client device may transmit a stream of audio data directly to the remote system in a synchronous manner and have the remote system generate remote gradients. In these examples, the stream of audio data may be discarded by the client device after generating the client gradients and / or by the remote system after generating the remote gradients, thereby preventing the stream of audio data from being stored in a non-transitory computer-readable storage medium of the client device or the remote system. In contrast, in response to determining that the connection status indicates weak or unstable between the client device and the remote system, the client device may implement federated learning techniques to generate client gradients locally at the client device and then transmit the client gradients to the remote system in an asynchronous manner once the client device establishes a strong or stable connection with the remote system.

[0009] In additional or alternative variations of these embodiments, the client device may make this determination based on the client device's current location (e.g., determined using a GPS sensor on the client device). For example, in response to determining that the client device's current location corresponds to a first location or a first geographic area, the client device may perform ephemeral learning techniques, generate client gradients locally on the client device, and synchronously transmit the client gradients to a remote system. Additionally or alternatively, the client device may transmit a stream of audio data directly to the remote system in a synchronous manner and have the remote system generate remote gradients. In these examples, the stream of audio data may be discarded by the client device after generating the client gradients and / or by the remote system after generating the remote gradients, thereby preventing the stream of audio data from being stored in a non-transitory computer-readable storage medium of the client device or the remote system. In contrast, in response to determining that the client device's current location corresponds to a second location or a second geographic region, the client device may implement federated learning techniques to generate client gradients locally at the client device, and then transmit the client gradients to the remote system in an asynchronous manner once the client device establishes a strong or stable connection with the remote system.

[0010] In various implementations, the client device and / or remote system may utilize various language identification techniques to identify a given language of a stream of speech utterances captured in a stream of audio data. In these implementations, the client device and / or remote system may perform ephemeral learning techniques and / or federated learning techniques only if it determines that the given language corresponds to a target language from among multiple target languages. Continuing with the above example in which a given user of a given client device is located in South Africa, further assume that the stream of audio data is capturing commercials from a given radio station in South Africa. The client device and / or remote system may use a language identification ML model to process the stream of audio data and identify a given language spoken in the given radio station's commercial. In this example, in response to determining that the given language spoken in the given radio station's commercial corresponds to English, the client device and / or remote system may determine that English is not one of the multiple target languages ​​(i.e., English is not considered a tail language) and refrain from performing ephemeral learning techniques and / or federated learning. In contrast, in response to determining that a given language spoken in a commercial on a given radio station corresponds to Swazi, the client device and / or remote system may determine that Swazi is one of multiple target languages ​​(i.e., Swazi may be considered a tail language) and perform ephemeral learning techniques and / or federated learning. Accordingly, the client device and / or remote system may utilize various language identification techniques to determine which streams of audio data capture streams of speech utterances in the given target language of interest when generating client and / or remote gradients for updating a global ML model for the given target language.

[0011] In various implementations, the client device and / or remote system may utilize various deduplication techniques and determine, based on the audio fingerprint, whether the stream of audio data has previously been used to generate client gradients and / or remote gradients for updating a global ML model for a given language. In these implementations, the client device and / or remote system may perform ephemeral learning techniques and / or federated learning techniques only if it determines that the stream of audio data has not previously been used to generate client gradients and / or remote gradients for updating a global ML model. For example, the client device and / or remote system may use an encoder-decoder ML model or other ML model to process the stream of audio data and generate an embedding (or another low-dimensional representation of the audio data) as an audio fingerprint for the stream of audio data. The embedding may be mapped to an embedding space (or other low-dimensional latent space), allowing the embedding to be compared to multiple embeddings previously generated for previously encountered streams of audio data. In these examples, if an embedding generated based on a stream of audio data matches a given embedding previously generated for a previously encountered given stream of audio data, ephemeral and / or federated training techniques are not performed and the stream of audio data may be discarded. In these examples, if the embedding and the previously generated given embedding are within a threshold distance in the embedding space (e.g., measured using Euclidean distance, cosine similarity, and / or other distance measures), ephemeral and / or federated training techniques are not performed and the stream of audio data may be discarded. Otherwise, the client device and / or remote system may perform ephemeral and / or federated training techniques to generate client gradients and / or remote gradients based on the stream of audio data.

[0012] Also, for example, the client device and / or remote system may process the stream of audio data using local sensitivity hashing and generate an audio hash as an audio fingerprint for the stream of audio data. In these examples, if the audio hash generated based on the stream of audio data matches a given audio hash previously generated for a previously encountered given stream of audio data (e.g., based on determining that the audio hash and the previously generated given audio hash match), the ephemeral learning techniques and / or federated learning techniques may not be performed and the stream of audio data may be discarded. Otherwise, the client device and / or remote system may perform ephemeral learning techniques and / or federated learning techniques to generate client gradients and / or remote gradients based on the stream of audio data.

[0013] In these embodiments, the client device (and other client devices) may transmit audio fingerprints to a remote system, enabling the remote system to generate a database of audio fingerprints. The database of audio fingerprints may be distributed from the remote system to the client device (and other client devices), enabling the client device (and other client devices) to employ these deduplication techniques. Thus, as a population of client devices generate these audio fingerprints, the remote system may maintain the database and periodically distribute the database of audio fingerprints to avoid duplicate and / or unnecessary processing on particular streams of audio data that the population of client devices encounters multiple times.

[0014] Continuing with the above example in which a given user of a given client device is located in South Africa, assume again that the stream of audio data captures commercials for a given radio station in South Africa. Further, assume that commercials for the given radio station in South Africa play every hour for a week. By using the deduplication techniques described herein, the client device and / or remote system prevents client gradients and / or remote gradients from being generated based on a stream of audio data capturing commercials every hour that the commercials play over the course of a week. Thus, these deduplication techniques not only save computational and / or network resources by avoiding performing ephemeral and / or federated learning techniques every time a commercial is encountered, but also prevent the global ML model from overfitting to the commercials, thereby resulting in a more robust global ML model in terms of accuracy and / or precision. Rather, these deduplication techniques enable the client device and / or remote system to limit the amount of instances in which the client device and / or remote system updates the global ML model based on client gradients and / or remote gradients generated based on multiple streams of audio data capturing the same stream of audio utterances.

[0015] In various implementations, the client device and / or the remote system may utilize unsupervised or self-supervised learning techniques, at least in part due to the lack of a teacher signal for the stream of audio data from the radio station. In some variations of these implementations, one non-limiting example of an unsupervised or self-supervised learning technique that may be utilized to generate the client gradients and / or the remote gradients corresponds to a teacher-learner technique. For example, the client device and / or the remote system may process the stream of audio data (e.g., using an on-device ML model and / or a global ML, respectively) to generate predicted output(s). Further, the predicted output(s) may be compared to corresponding benchmark output(s) generated based on benchmark ML model(s) that are also utilized to process the stream of audio data. In this case, the benchmark ML model(s) may be of the same type as the on-device ML model utilized by the client device and the global ML model utilized by the remote system, and the benchmark output(s) may be utilized as a quasi-teacher signal for generating the client gradients and / or the remote gradients, respectively. If the client device implements ephemeral learning techniques to generate client gradients based on processing using an on-device ML model, the benchmark ML model(s) may correspond to the on-device ML model, the global ML model, or a separate dedicated benchmark ML model. If the client device implements federated learning techniques to generate client gradients based on processing using an on-device ML model, the benchmark ML model(s) may correspond to the on-device ML model or a separate dedicated benchmark ML model. If the remote system implements ephemeral learning techniques to generate remote gradients based on processing using a global ML model, the benchmark ML model(s) may correspond to the on-device ML model, the global ML model, or a separate dedicated benchmark ML model.

[0016] In some further variations of these embodiments, the predicted output(s) may be utilized in generating client and / or remote gradients for updating the global ML model only upon determining that one or more conditions are met. The one or more conditions may include, for example, whether the predicted output(s) satisfy a predicted output threshold, whether the benchmark output(s) satisfy a benchmark output threshold, and / or other conditions. In other words, the predicted output(s) may be utilized in generating client and / or remote gradients for updating the global ML model only upon determining that the benchmark output(s) provide sufficient quasi-supervisory signals to update the global ML model.

[0017] In another variation of these embodiments, one non-limiting example of an unsupervised or self-supervised learning technique that can be utilized to generate client and / or remote gradients corresponds to a masking technique. For example, a target portion of a stream of audio data can be identified. The target portion of the stream of audio data can follow a preceding portion of the stream of audio data and precede a subsequent portion of the stream of audio data. Furthermore, the target portion of the stream of audio data can be masked using various techniques. The target portion of the stream of audio data can be arbitrarily selected or can be selected based on one or more criteria, such as a specific segment of n to m seconds of audio data corresponding to the target portion, and / or any other criteria for selecting the target portion of the client data. In this case, the target portion of the stream of audio data can correspond to a target audio waveform portion of the stream of audio data, the preceding portion of the stream of audio data can correspond to a preceding audio waveform portion received before the target audio waveform portion, and the subsequent portion of the stream of audio data can correspond to a subsequent audio waveform portion received after the target audio waveform portion.

[0018] In these implementations, a preceding portion of the stream of audio data and a following portion of the stream of audio data may be processed using an on-device ML model and / or a global ML model to generate predicted output(s) that predict a target portion of the stream of audio data. For example, in an implementation in which the target portion of client data corresponds to a target audio waveform portion of the stream of audio data, the preceding and following audio waveform portions may be processed using an on-device ML model and / or a global ML model to generate a predicted target audio waveform that is predicted to correspond to the target audio waveform portion. Further, the predicted target audio waveform may be compared to the masked target audio waveform to generate client and / or remote gradients. Thus, the on-device ML model may attempt to reconstruct the target audio waveform portion based on processing the preceding and following audio waveform portions. Notably, this technique may be particularly advantageous when the on-device ML model and the global ML model correspond to a multilingual ASR model because language may be irrelevant for reconstructing the target audio waveform portion.

[0019] In various implementations implementing ephemeral learning techniques, the stream of audio data may pass through buffer(s) (e.g., buffer(s) on the client device and / or buffer(s) on the remote system) before processing the stream of audio data. In some variations of these implementations, the buffer(s) may be utilized to assign various tags to the stream of audio data, the tags being based on a given language of the stream of audio utterances included in the stream of audio data (e.g., determined using language identification techniques described herein), whether the stream of audio data includes a stream of audio utterances that was previously used to update a global ML model (e.g., determined using deduplication techniques described herein), a name associated with the radio station (e.g., the name of an Internet radio station), an amplitude modulation (AM) radio band associated with the stream of audio data, a frequency modulation (FM) radio band associated with the stream of audio data, a current location of the client device when the stream of audio data was received, and / or a geographic region of the client device when the stream of audio data was received. In these implementations, if the client device and / or remote system determines that it is not necessary to process the stream of audio data, for example, when the stream of audio data includes a stream of audio utterances in a non-target language, when the stream of audio data includes a stream of audio utterances that was previously used to generate a gradient, and / or in other cases, the stream of audio data may be discarded from the buffer(s), and the client device and / or remote device system may refrain from performing any further processing on the stream of audio data.

[0020] As described herein, various architectures may be utilized to enable client devices and / or remote systems to receive streams of audio data. In some implementations, the streams of audio data may be generated by an Internet radio station. In some variations of these implementations, the streams of audio data may be received at the client device over one or more networks (e.g., the Internet) and processed locally at the client device using ephemeral and / or federated learning techniques to generate client gradients that are transmitted to the remote system. In additional or alternative variations of these implementations, the streams of audio data may be received at the client device over one or more networks (e.g., the Internet) and / or at other computing devices (e.g., one or more other servers) in communication with the remote system and transmitted to the remote system, causing the remote system to generate remote gradients. In additional or alternative implementations, the streams of audio data may be generated by an AM or FM radio station. In some variations of these embodiments, the stream of audio data may be received at the client device via an internal transceiver of the client device and / or via an external transceiver of the client device (e.g., an external transceiver that may be connected to the client device via a minijack on the client device or through one or more networks (e.g., Bluetooth)). In these embodiments, the client device may sweep AM and / or FM radio bands to receive the stream of audio data or may be tuned to a particular AM and / or FM radio band. The client device and / or remote system may implement ephemeral and / or federated learning techniques in the same or similar manner as described above to generate client and / or remote gradients.

[0021] In various embodiments, after updating the global ML model based on at least the client gradients and / or the remote gradients, the remote system may cause the updated global ML model (or updated global weights for the global ML model) to be distributed to the client device and / or additional client devices. In response to receiving the updated global ML model (or updated global weights for the global ML model), the client device and / or additional client devices may replace the on-device ML model with the updated global ML model (or replace the on-device weights for the on-device ML model with the updated global weights for the global ML model) in their on-device storage. In some variations of these embodiments, the remote system may cause the updated global ML model (or updated global weights for the global ML model) to be distributed to the client device and / or additional client devices only if it determines that one or more conditions are met. The one or more conditions may include, for example, whether the current time at the client device's current location corresponds to a particular time of day, whether the current day of the week at the client device's current location corresponds to a particular day of the week, whether the global ML model has been updated based on a threshold amount of gradients, and whether the performance of the updated global ML model meets a performance threshold and / or other criteria.

[0022] By using the techniques described herein, one or more technical advantages may be achieved. As one non-limiting example, the techniques described herein enable updating a global ML model based on processing streams of audio data available via radio stations that broadcast streams of audio data in different languages. By enabling updating a global ML model based on processing streams of audio data available via radio stations, it becomes possible to update the global ML model based on diverse streams of audio data in diverse languages ​​that may not otherwise be available. As a result, the global ML model may be more robust (e.g., in terms of precision and / or recall) and may be provided to a more diverse group of users. For example, the language identification techniques described herein enable updating a global ML model for particular tail languages ​​that are not well-defined or familiar to a majority of users. Also, for example, the deduplication techniques described herein enable updating a global ML model for diverse streams of audio data for these particular tail languages ​​that are not well-defined or familiar to a majority of users.

[0023] The above description is provided as a summary of only some of the embodiments disclosed herein. These and other embodiments are described in further detail herein. [Brief explanation of the drawings]

[0024] [Figure 1A] 1 illustrates exemplary process flows illustrating various aspects of the present disclosure that may be performed by client devices and remote systems to implement the embodiments disclosed herein. [Figure 1B] 1 illustrates exemplary process flows illustrating various aspects of the present disclosure that may be performed by client devices and remote systems to implement the embodiments disclosed herein. [Figure 2]FIG. 1 is a block diagram of an example environment in which various methods disclosed herein may be implemented. [Figure 3] A flowchart illustrating an exemplary method of ephemeral learning performed by a client device and based on a stream of audio produced by a given radio station is shown, according to various embodiments. [Figure 4] 1 is a flowchart illustrating an exemplary method of ephemeral learning performed by a remote system based on a stream of audio produced by a given radio station, according to various embodiments. [Figure 5] A flowchart illustrating an exemplary method for deduplication techniques utilized in ephemeral learning performed by a client device is shown, according to various embodiments. [Figure 6] A flowchart illustrating an exemplary method for deduplication techniques utilized in ephemeral learning performed by a remote system, according to various embodiments, is shown. [Figure 7] A flowchart illustrating an example method for determining whether a client device needs to perform ephemeral or federated learning techniques is presented, according to various embodiments. [Figure 8] 1 illustrates an exemplary architecture of a computing device, according to various embodiments. DETAILED DESCRIPTION OF THE INVENTION

[0025] 1A and 1B, an exemplary process flow illustrating various aspects of the present disclosure that may be implemented by client device 120 and remote system 160 is shown. Client device 120 shown in FIGS. 1A and 1B may include at least the components enclosed within the boxes representing client device 120 in FIGS. 1A and 1B. Similarly, remote system 150 shown in FIGS. 1A and 1B may include at least the components enclosed within the boxes representing remote system 150 in FIGS. 1A and 1B. As described herein, the process flows of FIGS. 1A and 1B may be utilized to perform ephemeral and / or federated training of audio-based machine learning (ML) model(s) based on streams of audio data generated via radio stations around the world.

[0026] 1A , assume that a stream of audio data 124 is received at a client device 120. The stream of audio data 124 may be received via an internal transceiver 122A of the client device 120 (e.g., integrated into the client device 120) or via an external transceiver 122B coupled to the client device (e.g., via an audio jack on the client device 120 or by other means). Furthermore, the stream of audio data 124 may originate from a variety of sources, including a radio station emitting a stream of audio data at different frequencies (e.g., amplitude modulation (AM) frequencies, frequency modulation (FM) frequencies, etc.) via a radio tower 110A, an Internet radio source 110B, and / or other radio sources. In some implementations, client device 120 may cause stream of audio data 124 to be buffered (e.g., in one or more buffers 132), while stream of audio data 124 is first processed by client device 120 to determine whether to perform ephemeral and / or federated training techniques on stream of audio data 124 and / or where to perform the ephemeral and / or federated training techniques (e.g., locally on client device 120 as shown in the process flow of FIG. 1A or on remote system 160 as shown in FIG. 1B).

[0027] As shown in FIG. 1A , client device 120 may include various engines for performing various operations. For example, client device 120 may include on-device ML model engine 134, gradient engine 136, learning engine 138, connection status engine 140, language identification engine 142, and deduplication engine 144. Additionally, remote system 160 may include remote update engine 180 and update distribution engine 182. While FIG. 1A depicts client device 120 and remote system 160 as including particular engines that perform process flows in particular ways, it should be understood that this is for purposes of illustrating the various techniques described herein and is not intended to be limiting. For example, the various engines shown in FIG. 1A may be combined, other engines may be added or omitted, and the process flows described with respect to FIG. 1A may be reordered (e.g., as described with respect to FIG. 1B ).

[0028] In general, the on-device ML model engine 134 may process the stream of audio data 124 using an on-device audio-based ML model to generate one or more predicted outputs 134A. The on-device audio-based ML model is the on-device counterpart of a global audio-based ML model (e.g., stored in global ML model(s) database 160A) that is stored and updated in on-device memory or storage (e.g., on-device ML model(s) database 120A) of the client device 120. The one or more predicted outputs 134A generated by the on-device audio-based ML model engine 134 may be based on the type of on-device audio-based ML model utilized in processing the stream of audio data 124 (and the type of global audio-based ML model updated based on processing the stream of audio data 124). For example, assume that the on-device audio-based ML model and the global audio-based ML model are corresponding language representation ML models. In this example, the one or more predicted outputs 134A may include, for example, a rich feature representation of the stream of audio data 124, including a description of one or more sounds captured in the stream of audio data 124, relationships between sounds captured in the stream of audio data 124, and / or other features of the stream of audio data 124. In contrast, the on-device audio-based ML model and the global audio-based ML model are assumed to be corresponding multilingual automatic speech recognition (ASR) models. In this example, the one or more predicted outputs 134A may include recognized text in a given language, such as, for example, one or more terms corresponding to a stream of speech utterances captured in the stream of audio data 124. While the above example is described with respect to a particular audio-based ML model, it should be understood that this is for illustrative purposes only and is not intended to be limiting, and that another non-limiting example of an audio-based ML model is described with reference to FIG. 2.

[0029] Additionally, the gradient engine 136 may generate gradients 136A based on the processing of the stream of audio data 124 by the on-device ML model engine 134 and / or one or more predicted outputs 134A generated by the on-device ML model engine 134. The gradients 136A may be transmitted to the remote system 160 for use in updating a global audio-based ML model, which is a global counterpart of the on-device audio-based ML model used in processing the stream of audio data 124. In generating the gradients 136A, the gradient engine 136 may utilize a learning engine 138. In some implementations, the learning engine 138 may employ one or more supervised learning techniques, while in other implementations, the learning engine 138 may employ one or more unsupervised or semi-supervised learning techniques. However, in various implementations, the stream of audio data 124 may be generated by a given radio station such that an explicit supervision signal for employing one or more of the supervised learning techniques is not available. Thus, although the techniques described herein are generally described with respect to one or more of unsupervised or semi-supervised learning techniques, one or more supervised learning techniques may additionally or alternatively be utilized by the learning engine 138.

[0030] In some variations of these embodiments, one or more of the unsupervised or semi-supervised learning techniques may correspond to a teacher-learner technique. In implementing a teacher-learner technique, the learning engine 138 may use one or more corresponding benchmark audio-based ML models to process the stream of audio data 124 and generate one or more benchmark outputs. In these embodiments, the one or more corresponding benchmark audio-based ML models may be the same ML model as the on-device audio-based ML model and / or the global audio-based ML model, or other audio-based ML models of the same type but different from the on-device audio-based ML model and the global audio-based ML model (e.g., models stored in the on-device ML model(s) database 120A and / or the global ML model(s) database 160A). Furthermore, one or more benchmark outputs may be used as quasi-teacher signals utilized in generating the gradients 136A. For example, the gradient engine 136 may compare one or more predicted outputs 134A to one or more benchmark outputs in generating the gradients 136A.

[0031] In some further variations of these embodiments, one or more benchmark outputs may be utilized as quasi-teacher signals only if it is determined that one or more conditions are met. The one or more conditions may include, for example, whether one or more of the predicted outputs 134A meet a predicted output threshold, whether one or more of the benchmark outputs meet a benchmark output threshold, and / or other conditions. In other words, one or more benchmark outputs may be utilized as quasi-teacher signals only if it is determined that the one or more benchmark outputs provide a valid teacher signal for one or more predicted outputs 134A.

[0032] For example, again assume that the on-device audio-based ML model and the global audio-based ML model are corresponding multilingual automatic speech recognition (ASR) models. In this example, the on-device ML engine 134 can process the stream of audio data 124 using the on-device multilingual ASR model and generate recognized text in a given language as one or more predicted outputs 134A. Additionally, the learning engine 138 can process the stream of audio data 124 using a benchmark multilingual ASR model and generate benchmark recognized text as one or more benchmark outputs. In this example, assuming one or more conditions are met, the gradient engine 136 can compare the recognized text and the benchmark recognized text in generating gradients 136A.

[0033] In additional or alternative variations of these embodiments, one or more of the unsupervised learning techniques or semi-supervised learning techniques may correspond to a masking technique. In implementing the masking technique, the learning engine 138 may identify a target portion of the stream of audio data 124 that is less than all of the audio data contained in the stream of audio data 124. Thus, the target portion of the stream of audio data 124 may be proximate to a preceding portion of the stream of audio data 124 that precedes the target portion of the stream of audio data 124 and / or a subsequent portion of the stream of audio data 124 that follows the target portion of the stream of audio data 124. The target portion of the stream of audio data 124 may be selected arbitrarily or may be selected based on one or more criteria, such as a specific segment of the stream of audio data 124 from n seconds to m seconds that has been identified as the target portion, and / or other criteria for selecting the target portion of the stream of audio data 124. Furthermore, the learning engine 138 may mask the target portion of the stream of audio data 124. Thus, the one or more streams of prediction outputs 134A may include one or more predictions for the masked target portions of the stream of audio data 124.

[0034] For example, again assume that the on-device audio-based ML model and the global audio-based ML model are corresponding language-representation ML models. In this example, the target portion of the stream of audio data 124 may correspond to a target audio waveform portion of the stream of audio data 124, the preceding portion of the stream of audio data 124 may correspond to a preceding audio waveform portion that precedes the target audio waveform portion, and the following portion of the stream of audio data 124 may correspond to a following audio waveform portion that follows the target audio waveform portion. Further, the on-device ML engine 134 can use the on-device language-representation ML model to process the preceding audio waveform portion that precedes the target audio waveform portion and / or the following audio waveform portion that follows the target audio waveform portion and generate one or more predicted outputs 134A, such as a prediction of the target audio waveform portion. Among other things, the prediction of the target audio waveform portion may include, for example, a predicted audio waveform of the target audio waveform, one or more predicted features of the target audio waveform (e.g., predicted amplitude, predicted wavelength, predicted phase, predicted period, and / or other features), one or more predicted features of the stream of audio data 124 (e.g., predicted MFCCs, predicted Melbank features, and / or other features), and / or other predicted representations of the target portion of the stream of audio data 124. In other words, the on-device ML model engine 134 may attempt to reconstruct the target audio waveform portion based on processing the preceding and / or subsequent audio waveform portions. In this example, the gradient engine 136 may compare the predictions made for the target portion of the stream of audio data 124 to the actual features of the stream of audio data 124 when generating gradients 136A.

[0035] While particular unsupervised or semi-supervised learning techniques are described herein, it should be understood that these techniques are provided for purposes of illustration and are not intended to be limiting. Rather, it should be understood that any unsupervised or semi-supervised learning technique that can be utilized in generating gradients based on processing the stream of audio data 124 can be utilized and is contemplated herein.

[0036] In some implementations, gradient 136A (and optionally one or more gradients 190A generated in the same or similar manner by one or more corresponding additional client devices 190, each including the same or similar components and / or engines described with respect to client device 120) may be generated in an ephemeral manner such that stream of audio data 124 is not stored in temporary memory or storage of client device 120. In some variations of these implementations, stream of audio data 124 may be discarded after generating gradient 136A. Additionally, gradient 136A may be transmitted to remote system 160 in a synchronous manner (e.g., in response to gradient 136A being generated). In other words, client device 120 may process stream 124 of audio data temporarily available at client device 120 to generate gradients 136A and synchronously transmit gradients 136A to remote system 160 (e.g., via one or more buffers 132) to reduce memory consumption at client device 120 (hence the term “ephemeral learning”). In additional or alternative implementations, client device 120 may update the on-device audio-based ML model utilized in generating gradients 136A and transmit one or more updated on-device weights of the updated on-device audio-based ML model to remote system 160 instead of gradients 136A.

[0037] In additional or alternative implementations, the gradient 136A (and optionally one or more gradients 190A generated in the same or similar manner by one or more corresponding additional client devices 190) may be generated in a federated manner, whereby the stream of audio data 124 may not be stored in temporary memory or storage of the client device 120. In some variations of these implementations, the stream of audio data 124 may be discarded after generation of the gradient 136A. However, the gradient 136A may be stored in memory or storage of the client device 120 and transmitted to the remote system 160 in an asynchronous manner (e.g., at a later time than when the gradient 136A was generated). In other words, the client device 120 may process the stream of audio data 124 to generate the gradient 136A, but may wait to transmit the gradient 136A to the remote system 160.

[0038] In these implementations, client device 120 may determine to perform ephemeral learning or federated learning, for example, based on the connection status of the connection between client device 120 and remote system 160 (and optionally, based on the connection strength of the connection). For example, connection status engine 140 may determine whether client device 120 is connected to remote system 160 (e.g., through one or more networks, such as one or more local area networks, one or more wide area networks, and / or one or more other networks) and / or may determine the connection strength between client device 120 and remote system 160 (e.g., through one or more of the networks). This determination may be made by connection status engine 140 while stream 124 of audio data is temporarily stored at client device 120 (e.g., in one or more of buffers 132). For example, in response to connection status engine 140 determining that the connection between client device 120 and remote system 160 through one or more of the networks is strong and / or stable, client device 120 may perform ephemeral learning and generate and send gradients 136A to remote system 160, because the connection status allows gradients 136A to be sent synchronously to remote system 160. In contrast, in response to connection status engine 140 determining that the connection between client device 120 and remote system 160 through one or more of the networks is weak and / or unstable, client device 120 may perform federated learning and generate and send gradients 136A to remote system 160. These techniques are described in more detail herein (e.g., with respect to FIG. 7 ).

[0039] In some implementations, client device 120 may process stream of audio data 124 to generate gradients 136A only if it determines that a given language of the stream of speech utterances captured in stream of audio data 124 corresponds to a target language. For example, language identification engine 142 may first process stream of audio data 124 to identify the given language using one or more language identification models (e.g., stored in on-device ML model(s) database 120A). Furthermore, language identification engine 142 may determine whether the given language corresponds to a target language (e.g., stored in target language(s) database 142A). Notably, the target language may be one of multiple target languages, e.g., defined by a developer, allowing the global audio-based ML model to be updated with respect to the target language. In other words, certain languages ​​(e.g., English, Spanish, French, German) may not be targeted due to the abundance of data available for updating the global audio-based ML model. These techniques therefore enable developers to target specific target languages ​​of interest (e.g., South African Swazi) for which there is little or no data to update global audio-based ML models. These techniques are described in more detail herein (e.g., with respect to FIG. 5).

[0040] In some implementations, client device 120 may process stream of audio data 124 to generate gradients 136A only if it determines that stream of audio data 124 has not previously been used to generate gradients for a global audio-based ML model update or has not previously been used to generate gradients for a global audio-based ML model update for more than a threshold amount of instances. For example, de-duplication engine 144 may first process stream of audio data 124 to generate an audio fingerprint for stream of audio data 124. Further, de-duplication engine 144 may determine whether the audio fingerprint matches a previously generated audio fingerprint (e.g., stored in audio fingerprint(s) database 144A). Among other things, the audio fingerprint may correspond to an embedding, an audio hash, and / or any other representation of stream of audio data 124 that enables stream of audio data 124 to be compared to another stream of audio data. In other words, the deduplication engine 144 may be utilized to prevent the global audio-based ML model from being continually updated based on the same stream of audio data (e.g., commercials on a given radio station), to prevent overfitting of the audio data, and to ensure diversity in the underlying stream of audio data utilized in generating the gradients. These techniques are described in more detail herein (e.g., with respect to FIG. 5).

[0041] In various implementations, assuming gradients 136A are transmitted to remote system 160 (e.g., in a synchronous or asynchronous manner), and assuming one or more gradients 190A are generated in the same or similar manner by one or more corresponding additional client devices 190, remote system 160 may update the global audio-based ML model using remote update engine 180 and then distribute updated audio-based ML model(s) 182A to client device 120 (and optionally one or more of the additional client devices 190) using update distribution engine 182. In some variations of these implementations, remote system 160 may update the global audio-based ML model in a streaming manner (e.g., as gradients 136A and one or more gradients 190A are received from each client device). In additional or alternative variations of these embodiments, remote system 160 may store gradients 136A and one or more gradients 190A in one or more databases (e.g., gradient(s) database 176B) and may update the global audio-based ML model in response to determining that one or more conditions for updating the global audio-based ML model are met. The one or more conditions for updating the global audio-based ML model may include, for example, a particular time of day, a particular day of the week, whether a threshold amount of gradients is available to update the global audio-based ML model, and / or other conditions.

[0042] When updating the global audio-based ML model, remote system 160 may update one or more weights of the global audio-based ML model based on gradients 136A and one or more gradients 190A received from each client device. In some implementations, remote update engine 180 can utilize a gradient descent algorithm to update one or more of the global weights. In some variations of these implementations, remote update engine 180 may average gradients 136A and one or more gradients 190A received from each client device before applying the gradient descent algorithm to update one or more of the global weights. In additional or alternative variations of these implementations, remote update engine 180 may utilize each, or a subset of, gradients 136A and one or more gradients 190A received from each client device to update one or more of the global weights using a gradient descent algorithm.

[0043] When distributing the updated audio-based ML model(s) 182A to client device 120 (and optionally one or more of additional client devices 190), update distribution engine 182 may determine whether one or more conditions (or one or more of its global weights) of updated audio-based ML model(s) 182A are satisfied. The one or more conditions may be based on whether client device 120 is ready to receive the updated audio-based ML model(s). For example, the one or more conditions may be based on whether client device 120 is charging, whether client device 120 has a charge state equal to or greater than a threshold, whether the temperature of client device 120 (based on one or more corresponding on-device temperature sensors) is below a threshold, whether client device 120 is being held by a user, temporal condition(s) associated with client device 120 (e.g., within a specific time period, every N hours (where N is a positive integer), and / or other temporal conditions), and / or other conditions. Furthermore, the one or more conditions may additionally or alternatively be based on another condition specific to remote system 160. For example, based on whether the performance of the updated audio-based ML model(s) 182A meets a performance threshold, whether the updated audio-based ML model(s) 182A are updated based on a threshold amount of gradient, etc., and / or other conditions.

[0044] Thus, when client device 120 (and optionally one or more of additional client devices 190) receives updated audio-based ML model(s) 182A (or one or more of its global weights), the on-device audio-based ML model (or one or more of its on-device weights) may be replaced with updated audio-based ML model(s) 182A (or one or more of its global weights). This process may be repeated to continue updating the global audio-based ML model. In some implementations, this process may be repeated to continue updating the global audio-based ML model until it is determined that the global audio-based ML model has converged.

[0045] 1B , assume again that stream of audio data 124 is received at client device 120. However, in FIG. 1B , in contrast to FIG. 1A , ephemeral learning is performed by remote system 160, and not performed locally at client device 120 and / or at one or more of additional client devices 190. In other words, in FIG. 1B , in contrast to FIG. 1A , client device 120 and / or one or more of additional client devices 190 are utilized as proxies for the remote system to obtain stream of audio data 124 from client device 120 and one or more streams of audio data 190B from a corresponding one of additional client devices 190.

[0046] 1B , remote system 160 includes remote-based counterparts of many of the components and engines described with respect to FIG. 1A . For example, one or more buffers 172 of the remote system of FIG. 1B may be utilized in the same or similar manner as described above with respect to one or more buffers 132 of client device 120. Global ML model engine 174 of remote system 160 may be utilized to generate one or more predicted outputs 174A in the same or similar manner as described with respect to on-device ML model engine 134, but using global audio-based ML models stored in global ML model(s) database 160A. Gradient engine 176 of remote system 160 may be utilized to generate gradients 176A at remote system 160 in the same or similar manner as described with respect to gradient engine 136 of the client device. The learning engine 178 of the remote system 160 may employ one or more supervised, unsupervised, or semi-supervised learning techniques in generating the gradients 176A in the same or similar manner as described with respect to the learning engine 138 of the client device 120. The language identification engine 184 of the remote system 160 may employ the same or similar manner as described with respect to the language identification engine 142 of the client device 120, but may employ remote memory or storage (e.g., a target language(s) database 184A) to determine whether a given language of the stream of speech utterances captured in the stream of audio data 124 corresponds to the target language. The deduplication engine 186 of the remote system 160 may employ the same or similar manner as described with respect to the deduplication engine 144 of the client device, but may employ remote memory or storage (e.g., an audio fingerprint(s) database 186A) to determine whether the stream of audio data 124 has previously been employed to update a global audio-based ML model.However, it should be noted that this determination may be made solely by client device 120, and therefore remote system 160 does not include a remote counterpart to connection status engine 140. Thus, it should be understood that the ephemeral learning described herein may be performed locally at client device 120, remotely from client device 120 (e.g., at remote system 160), or both.

[0047] Referring now to FIG. 2, a client device 250 is shown that, in this embodiment, includes various on-device ML engines as part of (or in communication with) the automated assistant client 240. Each ML model is also shown interfacing with a different on-device ML engine. Other components of the client device 250 are not shown in FIG. 2 for simplicity. FIG. 2 shows an example of how the various on-device ML engines and their respective ML models are utilized by the automated assistant client 240 in performing various actions.

[0048] 2 is shown with one or more microphones 211, one or more speakers 212, one or more other vision components 213, and display(s) 214 (e.g., a touch-sensitive display). Client device 250 may further include pressure sensor(s), proximity sensor(s), accelerometer(s), magnetometer(s), and / or other sensor(s) used to generate other sensor data in addition to the audio data captured by one or more microphones 211. Client device 250 selectively executes at least automated assistant client 240. Automated assistant client 240, in the example of FIG. 2, includes language expression engine 222, voice activity detection engine 224, hot word detection engine 226, ASR engine 228, multilingual ASR engine 230, continuous conversation engine 232, language identification engine 234, and speech identification engine 236. Automated assistant client 240 further includes audio capture engine 216 and visual capture engine 218. It should be understood that the ML engines and ML models shown in FIG. 2 are provided for illustrative purposes and are not intended to be limiting. For example, automated assistant client 240 may further include additional and / or alternative engines, such as a text-to-speech (TTS) engine and respective TTS models, an endpoint detector engine and respective endpoint detector models, and / or other engine(s) and associated ML model(s). It should further be understood that one or more of the engines and / or models described herein can be combined, such that a single engine and / or model can perform the functions of multiple engines and / or models described herein.

[0049] One or more cloud-based automated assistant components 270 can optionally be implemented on one or more computing systems (collectively referred to as "cloud" computing systems), which are communicatively connected to client device 250 via one or more networks generally designated as 299. Cloud-based automated assistant component 270 can be implemented, for example, via a cluster of high-performance servers. In various implementations, automated assistant client instance 240, through interaction with one or more cloud-based automated assistant components 270, can form what appears from a user's perspective as a logical instance of an automated assistant generally designated as 295, through which the user can engage in human-computer interaction (e.g., voice-based interaction, gesture-based interaction, and / or touch-based interaction).

[0050] Client device 250 may be a desktop computing device, a laptop computing device, a tablet computing device, a mobile phone computing device, a computing device in the user's vehicle (e.g., an in-car communication system, an in-car entertainment system, an in-car navigation system), a standalone interactive speaker, a smart appliance such as a smart television (or a standard television with a networked dongle with automated assistant functionality), and / or a user wearable device that includes a computing device (e.g., a user's watch with a computing device, a user's glasses with a computing device, a virtual reality or augmented reality computing device). Additional and / or alternative client devices may be provided.

[0051] The one or more vision components 213 may take various forms, such as a monographic camera, a stereographic camera, a LIDAR component (or other laser-based component(s)), a radar component, etc. The one or more vision components 213 may be used, for example, by a visual capture engine 218 to capture vision frames (e.g., image frames, laser-based vision frames) of the environment in which the client device 250 is deployed. In some implementations, such vision frame(s) may be utilized to determine whether a user is present near the client device 250 and / or the distance of a given user (e.g., the user's face) of the client device 250 relative to the client device 250. Such determination(s) may be utilized, for example, in determining whether to activate various on-device ML engines and / or other engine(s) shown in FIG. 2 . Additionally, the audio capture engine 218 may be configured to capture a user's voice utterance(s) and / or other audio data captured via one or more of the microphones 211.

[0052] As described herein, the stream of audio data is processed by the various engines shown in FIG. 2, and predictions can be made at client device 250 using corresponding ML models and / or at one or more of cloud-based automated assistant components 270 using corresponding ML models updated in the manner described herein (e.g., the manners with respect to FIGS. 1A, 1B, and 3-7).

[0053] As some non-limiting examples, each language expression engine 222, 272 can utilize its respective language expression model 222A, 272A to generate a rich feature representation of the stream of audio data and / or the stream of voice utterances captured in the stream of audio data. Each voice activity detection engine 224, 274 can utilize its respective voice activity detection model 224A, 274A to predict whether the stream of audio data contains voice activity of the user of the client device 250 and / or other users. Each hot word detection engine 226, 276 can utilize its respective language expression model 226A, 276A to predict whether the stream of audio data contains one or more specific words or phrases that invoke the automated assistant 295 (e.g., "OK Assistant," "Hey Assistant," "Assistant, what's the weather?", etc.) or specific functions of the automated assistant 295. Each ASR engine 228, 278 can utilize its respective ASR model 228A, 278A to generate recognized text for a given language based on processing a stream of audio data or can generate recognized text for a given language based thereon by predicting phoneme(s) and / or token(s) corresponding to a stream of audio data detected at the client device 250. Each multilingual ASR engine 230, 280 can utilize its respective ASR model 230A, 280A to generate recognized text in multiple languages ​​based on processing a stream of audio data or can generate recognized text for multiple languages ​​based thereon by predicting phoneme(s) and / or token(s) corresponding to a stream of audio data detected at the client device 250. Each continuous conversation engine 232, 282 can use its respective continuous conversation model 232A, 282A to predict whether an additional stream of audio data is intended for the automated assistant 295 (or, for example, for another user in the environment of the client device 250).Each language identification engine 234, 284 can utilize its respective language identification model 234A, 284A to predict a given language of a stream of vocal utterances captured in the stream of audio data, and each speech identification engine 236, 286 can utilize its respective speech identification model 236A, 286A to predict whether the stream of audio data captures a stream of vocal utterances of one or more users of client device 250 (e.g., by generating speaker embeddings or other representations and comparing them to actual embeddings corresponding to one or more of the users of client device 250).

[0054] In some implementations, one or more of the client device 250 and the cloud-based automated assistant component 270 may further include a natural language understanding (NLU) engine 238, 294 and a fulfillment engine 240, 296, respectively. The NLU engines 238, 294 may perform natural language understanding using the respective NLU models 238A, 294A on the recognized text, predicted phoneme(s), and / or predicted token(s) generated by the ASR engines 228, 278 and / or the multilingual ASR engines 230, 280 to generate NLU data. The NLU data may include, for example, intent(s) corresponding to the voice utterance and, optionally, slot value(s) for parameter(s) of the intent(s). Furthermore, one or more of the client device 250 and the cloud-based automated assistant component 270 may also include a fulfillment engine 240, 296, respectively. The fulfillment engines 240, 296 can generate fulfillment data based on processing the NLU data using their respective fulfillment models or rules 240A, 296A. This fulfillment data can define specific fulfillment in response to user input (e.g., voice utterances, typed input, touch input, gesture input, and / or any other type of user input) provided by a user of the client device 250. The specific fulfillment can include interaction(s) with locally installed application(s) based on the user input, command(s) based on the user input to send to Internet of Things (IoT) device(s) (directly or via corresponding remote system(s)), and / or other resolution action(s) based on the user input. The fulfillment data is then provided for local and / or remote performance / execution of the determined action(s) so that the specific fulfillment of the user input occurs.Execution may include, for example, rendering local and / or remote responses (e.g., visually and / or audibly (optionally utilizing an on-device TTS module)), interacting with locally installed applications, sending command(s) to IoT device(s), and / or other action(s). In another embodiment, the NLU engines 238, 294 and fulfillment engines 240, 296 may be omitted, and the ASR engines 228, 278 and / or the multilingual ASR engines 230, 280 may generate fulfillment data directly based on user input. For example, assume that the ASR engines 228, 278 and / or the multilingual ASR engines 230, 280 use their respective models to process a stream of audio data capturing the spoken utterance "Turn on the lights." In this example, the ASR engines 228, 278 and / or the multilingual ASR engines 230, 280 can generate a semantic output that is then sent to a software application associated with the light and / or sent directly to the light indicating that the light should be turned on.

[0055] Among other things, cloud-based automated assistant component(s) 270 include cloud-based counterparts to the engines and models described herein with respect to FIG. 2. However, in some implementations, the engines and models may not be utilized because they may be sent directly to client device 250 and executed locally on client device 250, while in other implementations, the engines and models may be utilized only when client device 250 detects any user input and sends the user input to cloud-based automated assistant component(s) 270. In various implementations, the engines and models executing on client device 250, cloud-based automated assistant component(s) 270 may be utilized in conjunction with each other in a distributed manner. Nevertheless, a remote execution module may optionally be included, which performs remote execution based on locally or remotely generated NLU data and / or fulfillment data. Additional and / or alternative remote engines may also be included. As described herein, in various implementations, on-device speech processing, on-device image processing, on-device NLU, on-device fulfillment, and / or on-device execution may be prioritized when resolving a voice utterance because they may at least reduce latency and / or network usage (because client-server round-trip(s) are not required to resolve the voice utterance). However, one or more cloud-based automated assistant component(s) 270 may be utilized, at least selectively. For example, such component(s) may be utilized in parallel with on-device component(s), and output from such component(s) may be utilized when processing of the local component(s) fails. For example, if processing of any of the on-device engines and / or models fails (e.g., due to relatively limited resources on the client device 250), the more robust resources of the cloud may be utilized.

[0056] Referring now to FIG. 3 , a flowchart illustrating an example method 300 of ephemeral learning performed by a client device and based on a stream of audio generated by a given radio station is shown. For convenience, the operations of method 300 are described with reference to a system on which the operations are performed. The system of method 300 includes one or more processors and / or other component(s) of a client device (e.g., client device 120 of FIGS. 1A and 1B , client device 250 of FIG. 2 , computing device 810 of FIG. 8 , and / or other client devices). Furthermore, while the operations of method 300 are shown in a particular order, this is not intended to be limiting. One or more operations may be reordered, omitted, or added.

[0057] In block 352, the system receives a stream of audio data capturing a stream of speech in a given language from a given radio station to which a user of the client device is actively listening. The stream of speech may capture, for example, speech containing commercial or advertising content for the given radio station, speech containing disc jockey content for the given radio station, speech containing podcast content for the given radio station, speech containing music content for the given radio station, and / or other content sources for the given radio station. However, in some implementations, the system may discard any stream of audio data capturing speech containing music content for the given radio station. In some implementations, the stream of audio data may be a fixed length (e.g., 5 seconds, 10 seconds, 15 seconds, etc.). In additional or alternative implementations, the stream of audio data may be a dynamic length (e.g., the length of commercial or advertising content, the length of disc jockey speech, the length of podcast content, etc.).

[0058] At block 354, the system generates gradients to update the global machine learning (ML) model for the given language. The system may return to block 352 and continue receiving additional streams of audio data capturing additional streams of voice utterances from the given radio station to which the client device user is actively listening. This allows the system to perform multiple iterations of method 300 of FIG. 3 in parallel, each time receiving a stream of audio data from the given radio station to which the client device user is actively listening. Notably, at block 354, no gradients have yet been generated. Rather, at block 354, the system may perform various operations (e.g., an audio fingerprinting operation described with respect to FIG. 5, a language identification operation described with respect to FIG. 5, and / or other operations described herein) to determine whether to generate gradients based on the stream of audio data.

[0059] In block 356, the system may determine whether to generate gradients locally on the client device or remotely off the client device. In some implementations, the system may determine to generate gradients locally on the client device by default. In other implementations, the system may determine to generate gradients remotely from the client device by default. In yet other implementations, the system may determine to generate gradients locally on the client device in some cases, but remotely off the client device in other circumstances. For example, the system may prioritize generating gradients locally on the client device to reduce network resource consumption when transmitting a stream of audio data to a remote system, but may begin transmitting the stream of audio data to the remote system in response to determining that a threshold amount of computational resources is being consumed on the client device. Also, for example, the system may periodically switch between generating gradients locally on the client device and remotely from the client device to more evenly distribute computational resources consumed by the client device and the remote system. These examples are not intended to be limiting, and it should be understood that any other criteria for determining whether to generate gradients locally on the client device or remotely off the client device are contemplated herein.

[0060] If, in the iteration of block 356, the system determines to generate gradients locally on the client device, the system may proceed to block 358. In block 358, the system processes the stream of audio data using an on-device ML model, which is an on-device counterpart of the global ML model and is stored in on-device memory or storage on the client device. In block 360, the system uses unsupervised or self-supervised learning techniques to generate gradients based on processing the stream of audio data using the on-device ML model. In block 362, the system discards the stream of audio data. In block 364, the system sends the gradients to a remote system and causes the remote system to update the global ML model based on the gradients. In other words, the system can cause the operations of the process flow described with respect to FIG. 1A to be performed locally on the client device.

[0061] If, in the iteration of block 356, the system determines to generate gradients remotely from the client device, the system may proceed to block 366. In block 366, the system sends the stream of audio data to a remote system and causes the remote system to update the global ML model based on processing the stream of audio data. In other words, rather than generating gradients locally on the client device based on processing the stream of audio data locally on the client device, the system may send the stream of audio data to a remote system and allow the remote system to process the stream of audio data and generate gradients based on processing the stream of audio data (e.g., as described with reference to FIG. 1B). In doing so, the system may cause the remote system to perform an action. For example, in block 366A, the system causes the remote system to process the stream of audio data using the global ML model. Further, in block 366B, the system causes the remote system to generate gradients based on processing the stream of audio data using unsupervised or self-supervised learning techniques. Further, in block 366C, the system causes the remote system to discard the stream of audio data. Further, in block 366D, the system causes the remote system to update the global ML model based on the gradients.

[0062] Thus, the system may first receive the stream of audio data at the client device and determine whether to generate gradients used to update the global ML model locally at the client device or remotely from the client device. In making this determination, the system considers the computational resources that would be consumed at the client device in generating the gradients and / or the network resources that would be consumed in transmitting the stream of audio data to the remote system, allowing the global ML model to be updated based on diverse streams of audio data generated by radio stations around the world while consuming fewer computational and / or network resources. As a result, the global ML model updated in this manner becomes more robust to processing and / or understanding more languages ​​of users around the world.

[0063] Referring now to FIG. 4 , a flowchart illustrating an example method 400 of ephemeral learning performed by a remote system based on a stream of audio generated by a given radio station is shown. For convenience, the operations of method 400 are described with reference to the system on which the operations are performed. The system of method 400 includes one or more processors and / or other component(s) of a remote system (e.g., remote system 160 of FIGS. 1A and 1B , cloud-based automated assistant component(s) 270 of FIG. 2 , computing device 810 of FIG. 8 , and / or other remote systems). Furthermore, while the operations of method 400 are shown in a particular order, this is not intended to be limiting. One or more operations may be reordered, omitted, or added.

[0064] At block 452, the system receives a stream of audio data from a given client device, the stream of audio data capturing a stream of speech in a given language. The stream of audio data was initially received at the given client device from a given radio station to which a user of the given client device actively listens. At block 454, the system processes the stream of audio data using a global ML model. The system may return to block 452 and continue receiving additional streams of audio data. At block 456, the system generates gradients using unsupervised or self-supervised learning techniques and based on processing the stream of audio data. At block 458, the system discards the stream of audio data. At block 460, the system updates the global ML model for the given language based on the gradients.

[0065] In other words, in an implementation according to method 400 of FIG. 4, the system may utilize a given client device as a proxy to obtain a stream of audio data. Further, the system may process the stream of audio data to generate gradients for use in updating a global ML model (e.g., as described with respect to FIG. 1B). This enables ephemeral learning to occur not only locally on the client device but also on remote systems. In additional or alternative implementations, the system may obtain a stream of audio data directly from a given radio station to further reduce consumption of network resources.

[0066] While method 400 of Figure 4 is described in a particular manner, it should be understood that this is for purposes of illustration and not limitation. For example, the system may perform multiple iterations of method 400 of Figure 4 in parallel for different streams of audio data received from a given client device and / or additional client devices. As another example, the system may first process a stream of audio data to determine whether to generate gradients to utilize in updating a global ML model (e.g., as described with respect to Figure 6).

[0067] Referring now to FIG. 5, a flowchart illustrating an example method 500 of a deduplication technique utilized in ephemeral learning implemented by a client device is shown. For convenience, the operations of method 500 are described with reference to a system that performs the operations. The system of method 500 includes one or more processors and / or other component(s) of a client device (e.g., client device 120 of FIGS. 1A and 1B, client device 250 of FIG. 2, computing device 810 of FIG. 8, and / or other client devices). Furthermore, while the operations of method 500 are shown in a particular order, this is not intended to be limiting. One or more operations may be reordered, omitted, or added.

[0068] In block 552, the system receives a stream of audio data capturing a stream of audio speech in a given language from a given radio station. The stream of audio speech may capture, for example, audio speech containing commercial or advertising content for the given radio station, audio speech containing disc jockey content for the given radio station, audio speech containing podcast content for the given radio station, audio speech containing music content for the given radio station, and / or other content sources for the given radio station. However, in some implementations, the system may discard any stream of audio data capturing audio speech containing music content for the given radio station. In some implementations, the stream of audio data may be a fixed length (e.g., 5 seconds, 10 seconds, 15 seconds, etc.). In additional or alternative implementations, the stream of audio data may be a dynamic length (e.g., the length of commercial or advertising content, the length of disc jockey speech, the length of podcast content, etc.). In some implementations, the stream of audio data may correspond to a stream of audio data that is actively listened to by a user of the client device when the user is actively listening to a given radio station. In additional or alternative implementations, the stream of audio data may correspond to a stream of audio data that is accessible at the client device via a given radio station but that is not actively listened to by a user of the client device.

[0069] At block 554, the system generates an audio fingerprint for the stream of audio data based on processing the stream of audio data. In some implementations, the audio fingerprint for the stream of audio data may correspond to an embedding. In some variations of these implementations, the system may process the stream of audio data using an encoder-decoder ML model or other ML model (e.g., one stored in on-device ML model(s) database 120A of FIGS. 1A and 1B) to generate an embedding (or another low-dimensional representation of the audio data). Further, the embedding may be mapped to an embedding space (or other low-dimensional latent space) that enables the embedding to be compared (e.g., as described with respect to block 556) with multiple embeddings previously generated for previously encountered streams of audio data (e.g., stored in audio fingerprint(s) database 144A of FIG. 1A). In additional or alternative implementations, the audio fingerprint for the stream of audio data may correspond to an audio hash. In some variations of these embodiments, the system processes the local sensitivity hash (e.g., stored in on-device ML model(s) database 120A of FIGS. 1A and 1B) to generate an audio hash (e.g., a vector representation of the stream of audio data), which in turn allows the audio hash to be compared (e.g., as described with respect to block 556) with multiple audio hashes previously generated for previously encountered streams of audio data (e.g., stored in audio fingerprint(s) database 144A of FIG. 1A).

[0070] At block 556, the system determines whether the stream of audio data has previously been used to generate gradients for updating a global ML model. For example, in embodiments in which the audio fingerprint corresponds to an embedding, if an embedding generated based on the stream of audio data matches a given embedding previously generated for a previously encountered given stream of audio data, the system may determine that the stream of audio data has previously been used to generate gradients for updating a global ML model. For example, if the embedding and the given embedding previously generated are within a threshold distance in the embedding space (e.g., measured using Euclidean distance, cosine similarity, and / or other distance measures), the system may determine that the embedding generated based on the stream of audio data matches the given embedding previously generated. Otherwise, the system may determine that the embedding generated based on the stream of audio data does not match any given embedding previously generated.

[0071] As another example, in an implementation in which an audio fingerprint corresponds to an audio hash, if an audio hash generated based on a stream of audio data matches a given audio hash previously generated for a given previously encountered stream of audio data, the system may determine that the stream of audio data was previously utilized to generate gradients for updating a global ML model. For example, the system may determine that an audio hash generated based on a stream of audio data matches a given previously generated audio hash if the corresponding vectors representing the stream of audio data satisfy a similarity threshold. Otherwise, the system may determine that an embedding generated based on a stream of audio data does not match any given previously generated embedding.

[0072] Among other things, in these embodiments, a remote system (e.g., including the client devices) communicatively coupled to a population of client devices may maintain and periodically distribute a database of audio fingerprints (e.g., audio fingerprint database 186A of FIG. 1B ). For example, each of the client devices in the population may generate an audio fingerprint generated locally on the client device and send it to the remote system, which may store the audio fingerprint in its database of audio fingerprints along with any audio fingerprints generated by the remote system. Additionally, the remote system may distribute the database of audio fingerprints (or updates thereto) to each of the client devices in the population to keep each database of audio fingerprints (e.g., audio fingerprint database 144A of FIG. 1A ) up to date. This prevents the population of client devices from continuously generating gradients based on the same stream of audio data and ensures sufficient diversity of the stream of audio data when generating gradients used to update the global ML model.

[0073] If, in the iteration of block 556, the system determines that the stream of audio data was previously used to generate gradients for updating the global ML model, the system may proceed to block 558. At block 558, the system refrains from further processing the stream of audio data. At block 560, the system discards the stream of audio data. The system may return to block 552 and perform additional iterations of method 500.

[0074] If, in the iteration of block 556, the system determines that the stream of audio data has previously been utilized to generate gradients for updating a global ML model, the system may proceed to block 562. In block 562, the system determines whether a given language corresponds to a target language. For example, the system may use a language identification model (e.g., stored in on-device ML model(s) database 120A) to process the stream of audio data (or recognized text (e.g., generated using a multilingual ASR model) generated based on processing the stream of audio data) to identify a given language associated with a stream of speech utterances captured in the stream of audio data. Furthermore, the system may compare the given language with one or more target languages ​​(e.g., stored in target language(s) database 142A of FIG. 1 ) to determine whether to generate gradients to be utilized for updating a global ML model for the given language. In other words, the system may discard streams of audio data capturing speech utterances in some languages ​​while continuing to process streams of audio data capturing speech utterances in other languages.

[0075] If, in the iteration of block 562, the system determines that the given language does not correspond to the target language, the system may proceed to block 558. In block 558, the system refrains from further processing the stream of audio data. In block 560, the system discards the stream of audio data. The system may return to block 552 and perform additional iterations of method 500. If, in the iteration of block 562, the system determines that the given language corresponds to the target language, the system may proceed to block 356 of method 300 of FIG. 3 and determine whether to perform ephemeral training locally on the client device or remotely from the client device. In other words, the system may have the client device perform these deduplication techniques to reduce the amount of computational and / or network resources consumed to process a previously processed stream of audio data. Furthermore, if the system determines that the gradients in question have not previously been used to generate gradients used to update a global ML model, the system may perform ephemeral training locally on the client device and remotely from the client device.

[0076] 5 is described in a particular manner, it should be understood that this is for purposes of illustration and not limitation. For example, the order of the de-duplication operations of blocks 556 and 562 may be reversed, such that the language identification de-duplication technique is performed before the audio fingerprint de-duplication technique. Also, for example, one of the de-duplication operations of blocks 556 and 562 may be omitted, such that only the language identification de-duplication technique or only the audio fingerprint de-duplication technique is performed.

[0077] Referring now to FIG. 6 , a flowchart illustrating an example method 600 of a deduplication technique utilized in ephemeral learning performed by a remote system is shown. For convenience, the operations of method 600 are described with reference to the system on which the operations are performed. The system of method 600 includes one or more processors and / or other component(s) of the remote system (e.g., remote system 160 of FIGS. 1A and 1B , cloud-based automated assistant component(s) 270 of FIG. 2 , computing device 810 of FIG. 8 , and / or other remote systems). Furthermore, while the operations of method 600 are shown in a particular order, this is not intended to be limiting. One or more operations may be reordered, omitted, or added.

[0078] At block 652, the system receives a stream of audio data from a given client device, the stream of audio data capturing a stream of speech in a given language, the stream of audio data originally received at the given client device from a given radio station. At block 654, the system generates an audio fingerprint for the stream of audio data based on processing the stream of audio data.

[0079] At block 656, the system determines whether the stream of audio data has previously been used to generate gradients for updating a global ML model. If, at the iteration of block 656, the system determines that the stream of audio data has previously been used to generate gradients for updating a global ML model, the system may proceed to block 658. At block 658, the system refrains from further processing the stream of audio data. At block 660, the system discards the stream of audio data. The system may return to block 652 and perform additional iterations of method 600. If, at the iteration of block 656, the system determines that the stream of audio data has previously been used to generate gradients for updating a global ML model, the system may proceed to block 662.

[0080] At block 662, the system determines whether the given language corresponds to the target language. If, at the iteration of block 662, the system determines that the given language does not correspond to the target language, the system may proceed to block 658. At block 658, the system refrains from further processing the stream of audio data. At block 660, the system discards the stream of audio data. The system may return to block 652 and perform additional iterations of method 600. If, at the iteration of block 662, the system determines that the given language corresponds to the target language, the system may proceed to block 456 of method 400 of FIG. 4 and utilize ephemeral learning at the remote system to generate gradients to update the global ML model.

[0081] In other words, method 600 of FIG. 6 is the same as or similar to method 500 of FIG. 5, but the system is implemented by a remote system in accordance with method 600 of FIG. 6 rather than a client device in accordance with method 500 of FIG. 5. Notably, in these implementations, the stream of audio data may be initially received by a given client device, but the operations of method 600 of FIG. 6 may be performed by components specific to the remote system. Furthermore, while method 600 of FIG. 6 is described with respect to a single client device and stream of audio data, it should be understood that this is for illustrative purposes and is not intended to be limiting. Rather, it should be understood that the system may cause multiple iterations of method 600 of FIG. 6 to be performed in parallel based on different streams of audio data received from different client devices.

[0082] Referring now to FIG. 7 , a flowchart illustrating an example method 700 for determining whether a client device should implement ephemeral or federated learning techniques is shown. For convenience, the operations of method 700 are described with reference to a system that performs the operations. The system of method 700 includes one or more processors and / or other component(s) of a client device (e.g., client device 120 of FIGS. 1A and 1B , client device 250 of FIG. 2 , computing device 810 of FIG. 8 , and / or other client devices). Furthermore, while the operations of method 700 are shown in a particular order, this is not intended to be limiting. One or more operations may be reordered, omitted, or added.

[0083] In block 752, the system receives a stream of audio data capturing a stream of audio speech in a given language from a given radio station (e.g., in the same or similar manner as described with respect to block 552 of method 500 of FIG. 5).

[0084] At block 754, the system determines whether to perform federated or ephemeral learning when generating gradients for updating the global ML model. For example, the system may determine whether to perform federated or ephemeral learning based on the connection status between the client device and the remote system that received the stream of audio data and / or the connection strength between the client device and the remote system (e.g., as described with respect to connection status engine 140 of FIG. 1A ). Additionally or alternatively, the system may determine whether to perform federated or ephemeral learning based on the location of the client device. For example, if the location of the client device is located within a particular geographic region, the system may determine to perform federated learning rather than ephemeral learning.

[0085] If, in the iteration of block 754, the system determines to perform federated learning, the system may proceed to block 756. In block 756, the system processes the stream of audio data using an on-device ML model, which is an on-device counterpart of the global ML model. In block 758, the system generates gradients based on processing the stream of audio data using unsupervised or self-supervised learning techniques and the on-device ML model. In block 760, the system asynchronously transmits the gradients to a remote system, causing the remote system to update the global ML model based on the gradients. The operations of blocks 756 and 758 may be performed in the same or similar manner as those described with respect to the operations of blocks 358 and 360, respectively, of method 300 of FIG. 3 .

[0086] However, in contrast to the operation of block 364 of method 300 of FIG. 3 , the transmission of gradients to the remote system at block 760 is asynchronous. In these embodiments, the system may have determined to perform federated learning to generate gradients based on a weak and / or unstable connection between the client device and the remote system. Thus, rather than attempting to transmit gradients to the remote system in response to having the gradients generated locally on the client device, the system may store the gradients in temporary memory or storage on the client device due to the weak and / or unstable connection between the client device and the remote system. However, the system may subsequently transmit the gradients to the remote system and have the remote system update the global ML model based on the gradients when a strong and / or stable connection is established between the client device and the remote system (e.g., over a Wi-Fi network, a cellular network, etc.). In some variations of these embodiments, the stream of audio data may be discarded in response to the gradients being generated, while in other variations of these embodiments, the system may store the stream of audio data in temporary memory or storage on the client device until the gradients are transmitted to the remote system in an asynchronous manner, and then discard the stream of audio data.

[0087] If, in the iteration of block 754, the system determines to perform ephemeral learning, the system may proceed to block 356 of method 300 of FIG. 3 . In other words, the system may have the client device determine whether to perform ephemeral learning or federated learning when generating gradients. Further, assuming the system determines to perform ephemeral learning based on a strong and / or stable connection between the client device and the remote system, the system may proceed to block 356 of method 300 of FIG. 3 to determine whether to generate gradients locally on the client device or remotely from the client device. In embodiments in which the system determines to perform ephemeral learning locally on the client device in the iteration of block 356, the system may, in the iteration of block 364 (e.g., in response to the gradients being generated), transmit the gradients to the remote system in a synchronous manner without storing them in temporary memory or storage on the client device. In embodiments in which the system determines to perform ephemeral learning remotely from the client device in an iteration of block 356, the system may transmit a stream of audio data to the remote system in a synchronous manner in an iteration of block 366.

[0088] It should be understood that while method 700 of Figure 7 and method 500 of Figure 5 are described as separate methods, this is for purposes of illustrating the techniques described herein and is not intended to be limiting. For example, the techniques described with respect to method 700 of Figure 7 and method 500 of Figure 5 (or some aspects thereof) may be combined into a single method.

[0089] 8, shown is a block diagram of an exemplary computing device 810 that may optionally be utilized to perform one or more aspects of the techniques described herein. In some implementations, one or more of the client device, cloud-based automated assistant component(s), and / or other component(s) may include one or more components of exemplary computing device 810.

[0090] Computing device 810 typically includes at least one processor 814 that communicates with several peripheral devices via a bus subsystem 812. These peripheral devices may include, for example, a storage subsystem 824 including a memory subsystem 825 and a file storage subsystem 826, user interface output devices 820, user interface input devices 822, and a network interface subsystem 816. The input and output devices enable user interaction with computing device 810. Network interface subsystem 816 provides an interface to an external network and is coupled to a corresponding interface device of another computing device.

[0091] The user interface input devices 822 may include a keyboard, a pointing device such as a mouse, trackball, touchpad, or graphics tablet, a scanner, a touchscreen integrated into a display, an audio input device such as a voice recognition system, a microphone, and / or other types of input devices. In general, use of the term "input device" is intended to include all possible types of devices and methods for inputting information into the computing device 810 or over a communications network.

[0092] The user interface output devices 820 may include a display subsystem, a printer, a fax machine, or a non-visual display such as an audio output device. The display subsystem may include a flat panel device such as a cathode ray tube (CRT), a liquid crystal display (LCD), a projection device, or some other mechanism for producing a visible image. The display subsystem may also provide a non-visual display, such as via an audio output device. In general, use of the term "output device" is intended to include all possible types of devices and methods for outputting information from the computing device 810 to a user or to another machine or computing device.

[0093] Storage subsystem 824 stores programming and data structures that provide the functionality of some or all of the modules described herein. For example, storage subsystem 824 may include logic for performing selected aspects of the methods disclosed herein or for implementing various components shown in Figures 1A, 1B, and 2.

[0094] These software modules generally execute on the processor 814 alone or in combination with other processors. The memory 825 used by the storage subsystem 824 may include several memories, including a main random access memory (RAM) 830 for storing instructions and data during program execution and a read-only memory (ROM) 832 in which fixed instructions are stored. The file storage subsystem 826 may provide persistent storage for program files and data files and may include a hard disk drive, a floppy disk drive with associated removable media, a CD-ROM drive, an optical drive, or a removable media cartridge. Modules that implement the functionality of particular embodiments may be stored by the file storage subsystem 826 in the storage subsystem 824, or may be stored on another machine accessible by the processor(s) 814.

[0095] The bus subsystem 812 provides a mechanism for allowing the various components and subsystems of the computing device 810 to communicate with each other as intended. Although the bus subsystem 812 is shown schematically as a single bus, alternative implementations of the bus subsystem may use multiple buses.

[0096] Computing device 810 can be of various types, including a workstation, a server, a computing cluster, a blade server, a server farm, or any other data processing system or computing device. Due to the ever-changing nature of computers and networks, the description of computing device 810 shown in Figure 8 is intended only as a specific example for purposes of describing some implementations. Many other configurations of computing device 810 can have more or fewer components than the computing device shown in Figure 8.

[0097] In situations where the systems described herein may collect or otherwise monitor personal information about users or utilize personal and / or monitored information, users may be provided with the opportunity to control whether a program or feature collects user information (e.g., information about the user's social network, social behavior or activities, occupation, user preferences, or the user's current geographic location) or whether and / or how it receives content from content servers that may be more relevant to the user. Certain data may also be processed in one or more ways to remove personally identifiable information before it is stored or used. For example, a user's identity may be processed so that information that can identify the user cannot be determined, or if geographic location information (such as to the city, zip code, or state level) is obtained, the user's geographic location may be generalized so that the user's specific geographic location cannot be determined. Thus, users may control how information about them is collected and / or used.

[0098] In some implementations, a method performed by one or more processors of a client device is provided, the method including receiving a stream of audio data capturing a stream of speech utterances in a given language from a given radio station; generating an audio fingerprint for the stream of audio data based on processing the stream of audio data; determining whether the stream of audio data has previously been utilized to generate gradients for updating a global machine learning (ML) model for the given language based on comparing the audio fingerprint for the stream of audio data with a database of audio fingerprints; in response to determining that the stream of audio data has not previously been utilized to generate gradients for updating a global ML model for the given language, processing the stream of audio data using an on-device ML model that is an on-device counterpart of the global ML model stored in on-device storage of the client device; generating gradients based on processing the stream of audio data using unsupervised or self-supervised learning techniques and using the on-device ML model; and transmitting the gradients to a remote system for utilization in updating the global ML model for the given language.

[0099] These and other implementations of the technology may include one or more of the following features.

[0100] In some implementations, the method may further include discarding the stream of audio data in response to determining that the stream of audio data was previously utilized in generating gradients for updating a global ML model for the given language.

[0101] In some embodiments, the method may further include receiving a database of audio fingerprints from the remote system and storing the database of audio fingerprints in on-device storage of the client device. In some variations of these embodiments, the remote system may have previously generated the database of audio fingerprints based on a plurality of corresponding streams of audio data received from the client device and a plurality of additional client devices capturing corresponding streams of audio utterances in a plurality of different languages, including the given language, from a plurality of different radio stations.

[0102] In some implementations, generating an audio fingerprint for the stream of audio data based on processing the stream of audio data may include processing the stream of audio data using local sensitivity hashes to generate an audio hash as the audio fingerprint. In some variations of these implementations, determining whether the stream of audio data has previously been utilized to generate a gradient for updating a global ML model for a given language based on comparing the audio fingerprint for the stream of audio data to a database of audio fingerprints may include comparing the audio hash generated based on processing the stream of audio data with a plurality of previously generated audio hashes, and determining whether the stream of audio data has previously been utilized to generate a gradient for updating a global ML model for the given language based on the comparison. The previously generated plurality of audio hashes may be previously generated based on processing a corresponding stream of audio data, and the previously generated plurality of audio hashes may be stored in the database of audio fingerprints.

[0103] In some implementations, generating an audio fingerprint for the stream of audio data based on processing the stream of audio data may include processing the stream of audio data using an encoder portion of an encoder-decoder ML model to generate the audio fingerprint as an embedding. In some variations of these implementations, determining whether the stream of audio data has previously been utilized to generate gradients for updating a global ML model for a given language based on comparing the audio fingerprint for the stream of audio data to a database of audio fingerprints may include comparing the embedding generated based on processing the stream of audio data with multiple previously generated embeddings, and determining based on the comparing whether the stream of audio data has previously been utilized to generate gradients for updating a global ML model for the given language. The multiple previously generated embeddings may be ones previously generated based on processing a corresponding stream of audio data, and the multiple previously generated embeddings may be stored in a database of audio fingerprints.

[0104] In some implementations, the method may further include processing the stream of audio data using an on-device language identification model stored in on-device storage of the client device to identify a given language and determining that the given language is one of a plurality of target languages. In some variations of these implementations, generating an audio fingerprint for the stream of audio data may be responsive to determining that the given language is one of the plurality of target languages. In some variations of these implementations, the method may further include refraining from further processing the stream of audio data and discarding the stream of audio data in response to determining that the given language is not one of the plurality of target languages. In some variations of these implementations, a developer associated with the global ML model may specify multiple target languages.

[0105] In some embodiments, the remote system may utilize the gradients to update one or more global weights of the global ML model to generate an updated global ML model. In some variations of these embodiments, the method may further include receiving one or more global weights of the updated global ML model from the remote system, or receiving the updated global ML model from the remote system. In some further variations of these embodiments, the method may further include replacing, in on-device storage of the client device, one or more on-device weights of the on-device ML model with one or more global weights of the updated global ML model, or replacing, in on-device storage of the client device, the on-device ML model with the updated global ML model.

[0106] In some implementations, the unsupervised or self-supervised learning techniques may include one or more of supervised learning techniques or masking techniques.

[0107] In some implementations, the global ML model may be a global feature extractor model that is updated to extract features from a stream of audio data for a given language.

[0108] In some implementations, the global ML model may be a multilingual automatic speech recognition (ASR) model that is updated to recognize text from a stream of audio data for a given language.

[0109] In some implementations, a method performed by one or more processors of a client device is provided, the method may include receiving a stream of audio data capturing a stream of speech in a given language from a given radio station, generating an audio fingerprint for the stream of audio data based on processing the stream of audio data, determining whether the stream of audio data has previously been utilized in generating gradients for updating a global machine learning (ML) model for the given language based on comparing the audio fingerprint for the stream of audio data with a database of audio fingerprints, and, in response to determining that the stream of audio data has not previously been utilized in generating gradients for updating the global ML model for the given language, transmitting the stream of audio data to a remote system. Sending the stream of audio data to the remote system causes the remote system to process the stream of audio data using the global ML model, generate gradients based on processing the stream of audio data using unsupervised or self-supervised learning techniques and with the global ML model, and update the global ML model based on the gradients.

[0110] These and other implementations of the technology may include one or more of the following features.

[0111] In some implementations, sending the stream of audio data to a remote system may further cause the remote system to discard the stream of audio data after generating the gradient.

[0112] In some embodiments, a method is provided that is performed by one or more processors of a remote system, the method including receiving, from a given client device, a stream of audio data capturing a stream of speech utterances in a given language, the stream of audio data being initially received at the given client device from a given radio station; generating an audio fingerprint for the stream of audio data based on processing the stream of audio data; determining, based on comparing the audio fingerprint for the stream of audio data with a database of audio fingerprints, whether the stream of audio data has previously been utilized in generating gradients for updating a global machine learning (ML) model for the given language; in response to determining that the stream of audio data has not previously been utilized in generating gradients for updating the global ML model for the given language, processing the stream of audio data using a global ML model; generating gradients based on processing the stream of audio data using unsupervised or self-supervised learning techniques and using the global ML model; and updating the global ML model for the given language based on the gradients.

[0113] In some implementations, a method performed by one or more processors of a client device is provided, the method comprising: receiving a stream of audio data capturing a stream of speech utterances in a given language from a given radio station; determining, based on a connection status between the client device and a remote system, whether to perform federated training or ephemeral training to generate gradients for updating a global machine learning (ML) model for the given language; and, in response to determining to perform federated training to generate gradients utilized to update the global ML model for the given language, processing the stream of audio data using an on-device ML model that is an on-device counterpart of the global ML model, stored in on-device storage of the client device. generating gradients based on processing the stream of audio data using unsupervised or self-supervised learning techniques and with the on-device ML model, and asynchronously transmitting the gradients to a remote system for use in updating a global ML model for the given language; and in response to determining to perform ephemeral learning to generate gradients for use in updating the global ML model for the given language, processing the stream of audio data using the on-device ML model, generating gradients based on processing the stream of audio data using unsupervised or self-supervised learning techniques and with the on-device ML model, and synchronously transmitting the gradients to a remote system for use in updating a global ML model for the given language.

[0114] These and other implementations of the technology may include one or more of the following features.

[0115] In some embodiments, the method may further include determining to perform federated learning to generate gradients for use in updating the global ML model for the given language based on a connection status between the client device and the remote system indicating that the client device is unable to connect to the remote system. In some variations of these embodiments, asynchronously sending the gradients to the remote system for use in updating the global ML model for the given language may include, after generating the gradients, determining that a connection has been established between the client device and the remote system, and, in response to determining that a connection has been established between the client device and the remote system, sending the gradients to the remote system for use in updating the global ML model for the given language.

[0116] In some embodiments, the method may further include determining, based on a connection status between the client device and the remote system indicating that the client device is connected to the remote system, to perform ephemeral learning to generate gradients for use in updating the global ML model for the given language. In some variations of these embodiments, synchronously transmitting the gradients to the remote system for use in updating the global ML model for the given language may include transmitting the gradients to the remote system for use in updating the global ML model for the given language without having to subsequently establish a connection between the client device and the remote system.

[0117] In some implementations, determining whether to perform federated or ephemeral learning to generate gradients for updating a global ML model for a given language may further be based on the location of the client device.

[0118] In some implementations, a method performed by one or more processors of a client device is provided, the method including receiving a stream of audio data capturing a stream of audio utterances in a given language from a given radio station; determining, based on a connection status between the client device and a remote system, whether to perform federated learning or ephemeral learning to generate gradients for updating a global machine learning (ML) model for the given language; in response to determining to perform federated learning to generate gradients for use in updating the global ML model for the given language, processing the stream of audio data using an on-device ML model that is an on-device counterpart of the global ML model stored in on-device storage of the client device; generating gradients based on processing the stream of audio data using unsupervised learning techniques and using the on-device ML model; asynchronously transmitting the gradients to the remote system for use in updating the global ML model for the given language; and in response to determining to perform ephemeral learning to generate gradients for use in updating the global ML model for the given language, synchronously transmitting the stream of audio data to the remote system. Synchronously transmitting the stream of audio data to the remote system causes the remote system to process the stream of audio data using a global ML model, generate gradients based on processing the stream of audio data using unsupervised or self-supervised learning techniques and with the global ML model, and update the global ML model based on the gradients.

[0119] These and other implementations of the technology may include one or more of the following features.

[0120] In some implementations, synchronously transmitting the stream of audio data to the remote system may further cause the remote system to discard the stream of audio data after generating the gradient.

[0121] In some implementations, a method is provided, performed by one or more processors of a client device, that includes receiving a stream of audio data capturing a stream of speech utterances in a given language from a given radio station actively listened to by a user of the client device, and generating gradients at a remote system for updating a global machine learning (ML) model for the given language. In some variations of these implementations, generating gradients includes processing the stream of audio data using an on-device ML model that is an on-device counterpart of the global ML model stored in on-device storage of the client device, generating gradients based on processing the stream of audio data using unsupervised or self-supervised learning techniques and using the on-device ML model, discarding the stream of audio data, and transmitting the gradients to the remote system for use in updating the global ML model for the given language. In another implementation, generating gradients includes transmitting the stream of audio data to the remote system. Sending the stream of audio data to the remote system causes the remote system to process the stream of audio data using a global ML model, generate gradients based on processing the stream of audio data using unsupervised or self-supervised learning techniques and with the global ML model, discard the stream of audio data, and update the global ML model based on the gradients.

[0122] In some implementations, a method performed by one or more processors of a remote system is provided, the method including receiving, from a given client device, a stream of audio data capturing a stream of audio utterances in a given language, the stream of audio data being initially received at the given client device from a given radio station to which a user of the given client device actively listens; processing the stream of audio data using a global machine learning (ML) model; generating gradients based on processing the stream of audio data using unsupervised or self-supervised learning techniques and using the global ML model; discarding the stream of audio data; and updating the global ML model for the given language based on the gradients.

[0123] Various embodiments may include a non-transitory computer-readable storage medium storing instructions executable by one or more processors (e.g., central processing unit (CPU)(ies), graphics processing unit (GPU)(ies), digital signal processor (DSP)(ies), and / or tensor processing unit (TPU)(ies)) to perform methods such as one or more of the methods described herein. Another embodiment may include an automated assistant client device (e.g., a client device including at least an automated assistant interface for interfacing with cloud-based automated assistant component(s)) including processor(s) operable to execute stored instructions to perform methods such as one or more of the methods described herein. Yet another embodiment may include a system of one or more servers including one or more processors operable to execute stored instructions to perform methods such as one or more of the methods described herein.

Claims

1. 1. A method implemented by one or more processors of a client device, comprising: receiving a stream of audio data from a given radio station, the stream capturing a stream of audio utterances in a given language; generating an audio fingerprint for the stream of audio data based on processing the stream of audio data; determining, based on comparing the audio fingerprint for the stream of audio data to a database of audio fingerprints, whether the stream of audio data has previously been utilized to generate gradients for updating a global machine learning (ML) model for the given language; in response to determining that the stream of audio data has not previously been utilized to generate gradients for updating the global ML model for the given language; processing the stream of audio data using an on-device ML model that is an on-device counterpart of the global ML model, stored in on-device storage of the client device; generating the gradients based on processing the stream of audio data using unsupervised or self-supervised learning techniques and using the on-device ML model; transmitting the gradients to a remote system for use in updating the global ML model for the given language; A method comprising:

2. in response to determining that the stream of audio data has previously been utilized to generate gradients for updating the global ML model for the given language; The method of claim 1 , further comprising discarding the stream of audio data.

3. receiving the database of audio fingerprints from the remote system; storing the database of audio fingerprints in the on-device storage of the client device; 3. The method of claim 1 or claim 2, further comprising:

4. 4. The method of claim 3, wherein the remote system previously generated the database of audio fingerprints based on a plurality of corresponding streams of audio data received from the client device and a plurality of additional client devices, capturing corresponding streams of audio utterances in a plurality of different languages, including the given language, from a plurality of different radio stations.

5. generating the audio fingerprint for the stream of audio data based on processing the stream of audio data, processing the stream of audio data using a local sensitivity hash to generate an audio hash as the audio fingerprint; 3. The method of claim 1 or claim 2, comprising:

6. determining whether the stream of audio data has previously been utilized to generate gradients for updating a global ML model for the given language based on comparing the audio fingerprint for the stream of audio data to a database of audio fingerprints; comparing the audio hash generated based on processing the stream of audio data with a plurality of previously generated audio hashes, the plurality of previously generated audio hashes having been previously generated based on processing a corresponding stream of audio data, the plurality of previously generated audio hashes being stored in the database of audio fingerprints; determining, based on said comparing, whether said stream of audio data has previously been utilized to generate gradients for updating a global ML model for said given language; The method of claim 5 , comprising:

7. generating the audio fingerprint for the stream of audio data based on processing the stream of audio data, processing the stream of audio data using an encoder portion of an encoder-decoder ML model to generate an embedding as the audio fingerprint; 3. The method of claim 1 or claim 2, comprising:

8. determining whether the stream of audio data has previously been utilized to generate gradients for updating a global ML model for the given language based on comparing the audio fingerprint for the stream of audio data to a database of audio fingerprints; comparing the embedding generated based on processing the stream of audio data with a plurality of previously generated embeddings, the previously generated embeddings having been previously generated based on processing a corresponding stream of audio data, the previously generated embeddings being stored in the database of audio fingerprints; determining, based on said comparing, whether said stream of audio data has previously been utilized to generate gradients for updating a global ML model for said given language; The method of claim 7, comprising:

9. processing the stream of audio data using an on-device language identification model stored in the on-device storage of the client device to identify the given language; determining whether the given language is one of a plurality of target languages, and wherein generating the audio fingerprint for the stream of audio data is responsive to determining that the given language is one of the plurality of target languages; The method of any one of claims 1 to 8, further comprising:

10. In response to determining that the given language is not one of the plurality of target languages, refraining from further processing the stream of audio data; and discarding the stream of audio data; 10. The method of claim 9, further comprising:

11. The method of claim 9 , wherein a developer associated with the global ML model specifies the multiple target languages.

12. The method of any one of claims 1 to 11, wherein the remote system utilizes the gradient to update one or more global weights of the global ML model to generate an updated global ML model.

13. receiving the one or more global weights of the updated global ML model from the remote system; or receiving the updated global ML model from the remote system; The method of claim 12 further comprising:

14. replacing one or more on-device weights of the on-device ML model with the one or more global weights of the updated global ML model in the on-device storage of the client device; or replacing the on-device ML model with the updated global ML model in the on-device storage of the client device; 14. The method of claim 13, further comprising:

15. The method of any one of claims 1 to 14, wherein the unsupervised or self-supervised learning techniques include one or more of teacher-learner or masking techniques.

16. The method of any one of claims 1 to 15, wherein the global ML model is a global feature extractor model that is updated to extract features from the stream of audio data for the given language.

17. The method of any one of claims 1 to 16, wherein the global ML model is a multilingual automatic speech recognition (ASR) model that is updated to recognize text from the stream of audio data for the given language.

18. 1. A method implemented by one or more processors of a client device, comprising: receiving a stream of audio data from a given radio station, the stream capturing a stream of audio utterances in a given language; generating an audio fingerprint for the stream of audio data based on processing the stream of audio data; determining, based on comparing the audio fingerprint for the stream of audio data to a database of audio fingerprints, whether the stream of audio data has previously been utilized to generate gradients for updating a global machine learning (ML) model for the given language; in response to determining that the stream of audio data has not previously been utilized to generate gradients for updating the global ML model for the given language; transmitting the stream of audio data to a remote system, wherein transmitting the stream of audio data to the remote system comprises: processing the stream of audio data using the global ML model; generating the gradients based on processing the stream of audio data using unsupervised or self-supervised learning techniques and using the global ML model; updating the global ML model based on the gradients; A method for making something happen.

19. Transmitting the stream of audio data to the remote system further comprises transmitting to the remote system: After generating the gradient, 20. The method of claim 18, further comprising causing the stream of audio data to be discarded.

20. 1. A method implemented by one or more processors of a remote system, comprising: receiving, from a given client device, a stream of audio data capturing a stream of audio speech in a given language, the stream of audio data having been originally received at the given client device from a given radio station; generating an audio fingerprint for the stream of audio data based on processing the stream of audio data; determining, based on comparing the audio fingerprint for the stream of audio data to a database of audio fingerprints, whether the stream of audio data has previously been utilized to generate gradients for updating a global machine learning (ML) model for the given language; in response to determining that the stream of audio data has not previously been utilized to generate gradients for updating the global ML model for the given language; processing the stream of audio data using the global ML model; generating the gradients based on processing the stream of audio data using unsupervised or self-supervised learning techniques and using the global ML model; updating the global ML model for the given language based on the gradients; and A method comprising:

21. 1. A method implemented by one or more processors of a client device, comprising: receiving a stream of audio data from a given radio station, the stream capturing a stream of audio utterances in a given language; determining whether to perform federated learning or ephemeral learning to generate gradients for updating a global machine learning (ML) model for the given language based on a connection status between the client device and a remote system; In response to determining to perform federated learning to generate the gradients utilized in updating the global ML model for the given language, processing the stream of audio data using an on-device ML model that is an on-device counterpart of the global ML model, stored in on-device storage of the client device; generating the gradients based on processing the stream of audio data using unsupervised or self-supervised learning techniques and using the on-device ML model; asynchronously transmitting the gradients to the remote system for use in updating the global ML model for the given language; In response to determining to perform ephemeral learning to generate the gradients utilized in updating the global ML model for the given language, processing the stream of audio data using the on-device ML model; generating the gradients based on processing the stream of audio data using the unsupervised or self-supervised learning technique and using the on-device ML model; synchronously transmitting the gradients to the remote system for use in updating the global ML model for the given language; A method comprising:

22. 22. The method of claim 21 , further comprising: determining to perform federated learning to generate the gradients utilized in updating the global ML model for the given language based on the connection status between the client device and the remote system indicating that the client device is unable to connect to the remote system.

23. asynchronously transmitting the gradients to the remote system for use in updating the global ML model for the given language, After generating the gradient, determining that a connection is established between the client device and the remote system; In response to determining that the connection is established between the client device and the remote system, transmitting the gradients to the remote system for use in updating the global ML model for the given language; 23. The method of claim 22, comprising:

24. 24. The method of claim 21, further comprising: determining, based on the connection status between the client device and the remote system indicating that the client device is connected to the remote system, to perform ephemeral learning to generate the gradients utilized in updating the global ML model for the given language.

25. synchronously transmitting the gradients to the remote system for use in updating the global ML model for the given language, transmitting the gradients to the remote system for use in updating the global ML model for the given language without the need to subsequently establish a connection between the client device and the remote system; 25. The method of claim 24, comprising:

26. 26. The method of claim 21, wherein determining whether to perform federated or ephemeral learning to generate gradients for updating the global ML model for the given language is further based on a location of the client device.

27. 1. A method implemented by one or more processors of a client device, comprising: receiving a stream of audio data from a given radio station, the stream capturing a stream of audio utterances in a given language; determining whether to perform federated learning or ephemeral learning to generate gradients for updating a global machine learning (ML) model for the given language based on a connection status between the client device and a remote system; In response to determining to perform federated learning to generate the gradients utilized in updating the global ML model for the given language, processing the stream of audio data using an on-device ML model that is an on-device counterpart of the global ML model, stored in on-device storage of the client device; generating the gradients based on processing the stream of audio data using unsupervised or self-supervised learning techniques and using the on-device ML model; asynchronously transmitting the gradients to the remote system for use in updating the global ML model for the given language; In response to determining to perform ephemeral learning to generate the gradients utilized in updating the global ML model for the given language, and synchronously transmitting the stream of audio data to the remote system, wherein synchronously transmitting the stream of audio data to the remote system comprises: processing the stream of audio data using the global ML model; generating the gradients based on processing the stream of audio data using the unsupervised or self-supervised learning technique and using the global ML model; updating the global ML model based on the gradients; A method for making something happen.

28. Synchronously transmitting the stream of audio data to the remote system further comprises, on the remote system:

28. The method of claim 27, further comprising discarding the stream of audio data after generating the gradient.

29. 1. A method implemented by one or more processors of a client device, comprising: receiving a stream of audio data capturing a stream of speech utterances in a given language from a given radio station that is actively listened to by a user of the client device; generating gradients at a remote system to update a global machine learning (ML) model for the given language; generating the gradient comprises: processing the stream of audio data using an on-device ML model that is an on-device counterpart of the global ML model, stored in on-device storage of the client device; generating the gradients based on processing the stream of audio data using unsupervised or self-supervised learning techniques and using the on-device ML model; discarding the stream of audio data; transmitting the gradients to the remote system for use in updating the global ML model for the given language; or generating the gradient comprises: transmitting the stream of audio data to the remote system, wherein transmitting the stream of audio data to the remote system includes: processing the stream of audio data using the global ML model; generating the gradients based on processing the stream of audio data using the unsupervised or self-supervised learning technique and using the global ML model; discarding the stream of audio data; updating the global ML model based on the gradients; A method for making something happen.

30. 1. A method implemented by one or more processors of a remote system, comprising: receiving, from a given client device, a stream of audio data capturing a stream of audio speech in a given language, the stream of audio data originally received at the given client device from a given radio station to which a user of the given client device is actively listening; processing the stream of audio data using a global machine learning (ML) model; generating gradients based on processing the stream of audio data using unsupervised or self-supervised learning techniques and using the global ML model; discarding the stream of audio data; updating the global ML model for the given language based on the gradients; and A method comprising:

Citation Information

Patent Citations

  • Sorting prediction device and storage medium storing computer program

    JP1999096132A

  • Dialog system, dialog device, response controller, control method of dialog device, control method of response controller, and control program

    JP2018185431A