Use corrections from automated assistant features to train on-device machine learning models
By using local machine learning models to process audio and sensor data on client devices, deciding whether to activate the automation assistant function and using gradient update model weights, the false negative and false positive problems in the prior art are solved, and more efficient and accurate automation assistant function is achieved.
Patent Information
- Application Number
- CN201980101834.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-11-08
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2039-11-08
AI Technical Summary
When existing automation assistants determine whether to activate the function, they are prone to false negatives and false positives, resulting in waste of resources and leakage of user privacy.
By processing audio data and other sensor data using locally stored machine learning models on the client device, predicted output is generated and decisions are made to activate the automation assistant function based on the predicted output. At the same time, the gradient is used to update the weight of the machine learning model and improve the model performance.
Reduces the frequency of false negatives and false positives, improves the accuracy and efficiency of the automation assistant function, protects user privacy and saves resources.
Smart Images

Figure CN114651228B_ABST
Abstract
Description
Background Art
[0001] Humans may engage in human-computer conversations using interactive software applications, referred to herein as "automated assistants" (also referred to as "digital agents," "interactive personal assistants," "intelligent personal assistants," "assistant applications," "conversational agents," and the like). For example, humans (who may be referred to as "users" when they interact with automated assistants) may provide commands and / or requests to automated assistants by providing textual (e.g., typed) natural language input and / or through physical movements (e.g., gestures, eye gaze, facial movements, and the like) without touch and / or speech, using spoken natural language input (i.e., utterances), which in some cases may be converted to text and then processed. The automated assistant responds to the request by providing responsive user interface output (e.g., audible and / or visual user interface output), controlling one or more smart devices, and / or controlling one or more functions of a device implementing the automated assistant (e.g., controlling other applications of the device).
[0002] As described above, many automated assistants are configured to interact via spoken utterances. To protect user privacy and / or to conserve resources, the automated assistant refrains from performing one or more automated assistant functions based on all spoken utterances present in audio data detected via a microphone of a client device that (at least in part) implements the automated assistant. Conversely, certain processing based on spoken utterances occurs only in response to determining that certain conditions exist.
[0003] For example, many client devices that include automated assistants and / or interface with automated assistants include a hot word detection model. When the microphone of such a client device is not deactivated, the client device can use the hot word detection model to continuously process audio data detected via the microphone to generate a prediction output indicating whether there are one or more hot words (including multi-word phrases), such as "Hey Assistant", "OK Assistant" and / or "Assistant". When the prediction output indicates the presence of a hot word, any audio data that follows within a threshold amount of time (and optionally, is determined to include voice activity) can be processed by one or more devices and / or remote automated assistant components (such as speech recognition components, sound activity detection components, etc.). In addition, a natural language understanding engine can be used to process the recognized text (from the speech recognition component), and / or actions can be performed based on the natural language understanding engine output. The action may include, for example, generating and providing a response and / or controlling one or more applications and / or smart devices. However, when the prediction output indicates that there are no hot words, the corresponding audio data will be discarded without any further processing, thereby saving resources and protecting user privacy.
[0004] Some automated assistants additionally or alternatively implement a continue conversation mode that can be enabled. When enabled, the continue conversation mode can process any spoken input detected via a microphone of a client device within a threshold amount of time of a previous spoken utterance directed to the automated assistant and / or within a threshold amount of time after the automated assistant has performed an action based on the previous spoken utterance. For example, a user may initially invoke the automated assistant (e.g., via a hotword, a hardware or software button, etc.) and provide an initial utterance of "Turn on the living room lights," followed by a subsequent utterance of "Turn on the kitchen lights." When the continue conversation mode is enabled, the automated assistant will act on the subsequent utterance without the user invoking the assistant again.
[0005] The continued conversation mode can distinguish between subsequent utterances of a user that are intended to be processed by the automated assistant and utterances that are not so intended (e.g., utterances that are intended for another person instead). In doing so, a machine learning model can be optionally used to process audio data that captures the subsequent utterance together with recognized text and / or representations thereof from the subsequent utterance (e.g., natural language understanding data generated based on the recognized text). A prediction output is generated based on the processing and indicates whether the subsequent utterance is intended for the automated assistant. Another automated assistant function is activated only if the prediction output indicates that the subsequent utterance is intended for the automated assistant. Otherwise, the other automated assistant function is not activated and the data corresponding to the subsequent utterance is discarded. Another function can include, for example, another verification that the subsequent utterance is intended for the automated assistant and / or an action is performed based on the subsequent utterance.
[0006] The above and / or other machine learning models (e.g., the additional machine learning models described below) perform well in many cases, with their predicted outputs specifying whether the automated assistant function is activated. However, there are still occurrences of false negative determinations and false positive determinations based on machine learning models.
[0007] With false negatives, a prediction output specifies that automated assistant features are not activated, even though the audio data (and / or other data) processed to generate the prediction output is suitable for activating those features. For example, assume that the prediction output generated using the hotword detection model is a probability, and that the probability must be greater than 0.85 before the automated assistant feature is activated. If the spoken utterance does include a hotword, but the prediction output generated based on processing the audio data is only 0.82, the feature will not be activated and this will be considered a false negative. The occurrence of false negatives can prolong the human / automated assistant interaction, forcing the human to repeat the utterance (and / or perform other actions) originally intended to activate the automated assistant feature.
[0008] With a false positive, a prediction output specifies activation of automated assistant features, even though the audio data (and / or other sensor data) that was processed to generate the prediction output is not appropriate for activating those features. For example, assume that the prediction output generated using a hotword detection model is a probability, and that probability must be greater than 0.85 before the automated assistant feature is activated. If the spoken utterance does not include a hotword, but the prediction output generated based on processing the audio data is 0.86, the feature will still be activated and this will be considered a false positive. In addition to privacy considerations, the occurrence of false positives can waste network and / or computing resources by activating features unnecessarily. Summary of the invention
[0009] Some implementations disclosed herein relate to improving the performance of the machine learning model utilized when determining whether to initiate an automated assistant function. As described in more detail herein, such a machine learning model may include, for example, a hot word detection model, a continued conversation model, a hot word call model, and / or other machine learning models. Various implementations generate prediction outputs at a client device based on processing audio data and / or other sensor data using a machine learning model stored locally at a client device. Those implementations also make decisions about whether to initiate one or more automated assistant functions based on the prediction output. For example, the decision can be based on whether the prediction output meets a threshold. In addition, those implementations determine whether the decision made based on the prediction output is correct locally at the client device and based on analyzing another user interface input and / or other data. When it is determined that the decision is incorrect (i.e., the decision is a false negative or a false positive), those implementations generate gradients locally at the client device based on comparing the prediction output with a ground truth output (e.g., a ground truth output that meets a threshold).
[0010] In some implementations, the generated gradients are used by one or more processors of the client device to update one or more weights of the machine learning model based on the generated gradients. For example, back propagation and / or other techniques may be used to update the weights based on the gradients. This may improve the performance of the machine learning model stored locally at the client device, thereby mitigating the occurrence of false negatives and / or false positives based on the predicted output generated using the machine learning model. In addition, this enables the performance of the machine learning model on the device to be improved for the attributes of the user of the client device (such as pitch, intonation, accent, and / or other speech characteristics in the case of a machine learning model processing audio data that captures spoken utterances).
[0011] In some implementations, the generated gradients are additionally or alternatively transmitted by the client device to the remote system via the network. In those implementations, the remote system updates the global weights of the corresponding global machine learning model using the generated gradients and the additional gradients from the additional client device. Based on determining that the corresponding decision is incorrect, additional gradients from the additional client device can be generated similarly locally at the corresponding additional client device. In various implementations, the client device transmits the generated gradients without transmitting any data (e.g., audio data and / or other sensor data) used to generate the predicted output determined to be incorrect, and does not transmit any data (e.g., another user interface input) used to determine that the predicted output is incorrect. The remote system can utilize the generated gradients when updating the global model without reference or use of such data. Compared with transmitting a larger data size data used to generate the predicted output and determine that the predicted output is incorrect, transmitting only the gradients utilizes fewer network resources. In addition, the transmission of the gradients protects the privacy and security of personal data because the data used when generating the predicted output and determining that the predicted output is incorrect cannot be derived from the gradients. In some implementations, one or more differential privacy techniques (e.g., adding Gaussian noise) can be used to further ensure that such data cannot be derived from the gradients.
[0012] In implementations where the remote system updates the global weights of the speech recognition model, the remote system may thereafter provide the updated global weights to the client device so that the client device replaces the weights of the machine learning model on its device with the updated global weights. In some implementations, the remote system may additionally or alternatively provide the updated machine learning model to the client device so that the client device replaces the machine learning model on their device with the updated global machine learning model. Thus, performance on the device is improved by utilizing the updated global weights or the updated global machine learning model.
[0013] Various techniques can be used to determine whether the decision to initiate the currently dormant automated assistant function is incorrect. In many implementations, determining that the decision is incorrect can be based on another user interface input received at the client device after the sensor data used to make the decision and contradicting the decision (explicitly or implicitly). In those implementations, determining that the decision is incorrect can be based on the duration between receiving the sensor data used to make the decision and receiving another user interface input. For example, the probability of determining that the decision is incorrect can increase as the duration decreases, and / or the probability of determining that the decision is incorrect can depend on the duration being less than a threshold. In those implementations, determining that the decision is incorrect can be additionally or alternatively based on a determined measure of similarity between another user interface input and the sensor data used to make the decision (wherein the probability of determining that the decision is incorrect increases as the similarity indicated by the measure of similarity increases). For example, the measure of similarity can be based on a duration similarity based on a comparison of the duration of another user interface input with the duration of the sensor data used to make the determination. In addition, for example, when the other user interface input is an additional spoken utterance and the sensor data used to make the determination includes a previous spoken utterance, the measure of similarity can be based on acoustic similarity based on a comparison of acoustic characteristics of the spoken utterance with acoustic characteristics of the additional spoken utterance and / or based on textual similarity based on a comparison of recognized text of the spoken utterance with recognized text of the additional spoken utterance.
[0014] In some implementations, determining whether a decision is incorrect can be based on the magnitude of a predicted output generated by a corresponding machine learning model and used to make a decision. In some of those implementations, the decision on whether to initiate a currently dormant automated assistant function can depend on whether the magnitude of the predicted output meets a threshold, and whether that decision is determined to be incorrect can be based on how close the predicted output is to the threshold. For example, assume that the predicted output indicates a probability, and in order to initiate the automated assistant function, the probability must be greater than 0.85. In such an example, determining whether a decision not to initiate an automated assistant function is incorrect can be based on how close the probability is to the threshold. For example, the closer the probability is to the threshold, the more likely the decision is to be determined to be incorrect and / or can depend on the probability being within a specific range of the threshold. Considering the magnitude of the predicted output can prevent incorrectly determining that a true negative is a false negative and / or determining that a true positive is a false positive.
[0015] Some specific examples of determining whether a decision about whether to initiate a currently dormant automated assistant function is incorrect are now provided with reference to a hotword detection model that monitors for the presence of an invocation hotword that, when detected, initiates certain processing of audio data that follows within a threshold amount of time of the invocation hotword.
[0016] As an example, assume that the hot word detection model is trained to generate a prediction output indicating the probability of whether the hot word "good assistant" is present in the audio data, and if the probability is greater than 0.85, the hot word is determined to exist. It is also assumed that the initial spoken utterance captured in the audio data includes the hot word "good assistant", but the prediction output generated based on the processing of the audio data indicates a probability of only 0.8. It is further assumed that the subsequent spoken utterance captured in the additional audio data is received 2.0 seconds after the initial spoken utterance includes the hot word "good assistant" (e.g., after the initial spoken utterance is completed), and the prediction output generated based on the processing of the additional audio data indicates a probability of 0.9. Therefore, in this example, the user initially says "good assistant" to call the assistant, which is not recognized as calling the hot word, and the user quickly follows another instance of "good assistant" to try to call the assistant again—and the following instance is recognized as calling the hot word. In this example, based on 0.8 being less than 0.85, an initial decision will be made to respond to the initial spoken utterance without initiating certain processing of the audio data. However, a subsequent decision to initiate certain processing of the audio data in response to the subsequent spoken utterance will be made based on 0.9 being greater than 0.85. In addition, in this example, it can be determined that the initial decision is incorrect. This can be based on another user interface input (i.e., a subsequent spoken utterance) that meets the threshold, can be based on the probability of the initial spoken utterance (0.8), the duration (2.0 seconds) between receiving the initial spoken utterance and the subsequent spoken utterance (which contradicts the first spoken utterance by meeting the threshold), and / or can be based on determining that the initial spoken utterance and the subsequent spoken utterance may be from the same user (e.g., using a speaker identification). For example, determining that the previous decision is incorrect can depend on the duration being less than a threshold duration (e.g., 4.0 seconds or other threshold duration) and / or on the probability of the initial decision being within a range of a 0.85 threshold (e.g., within a range of 0.35 or other ranges). In other words, in this case, an incorrect decision is determined only when the duration is less than the threshold duration and the probability is within a threshold range of the 0.85 threshold.
[0017] In addition, for example, determining that a previous decision is incorrect can additionally or alternatively be a function of the duration and probability of the initial decision, optionally also not necessarily dependent on satisfying an individual threshold. For example, the difference between the probability and the 0.85 threshold can be determined and multiplied by a factor based on duration, and the resulting value can be compared to the threshold to determine whether the decision is incorrect. For example, if the resulting value is less than 0.25, the resulting value can indicate a correction, and if the duration is 0.0 to 1.5 seconds, the duration-based factor can be a factor of 0.5, if the duration is 1.5 seconds to 3.0 seconds, the duration-based factor can be a factor of 0.6, if the duration is 3.0 seconds to 6.0 seconds, the duration-based factor can be a factor of 1.0, and if the duration is greater than 6.0 seconds, the duration-based factor can be a factor of 8.0. Therefore, in the case of this example, the difference (0.05) can be multiplied by 0.6 (corresponding to a factor of 2.0 seconds duration) to determine a value of 0.03 that is less than 0.25. Compare it with an alternative example with the same difference (0.05) but a duration of 7.0 seconds. In such an alternative example, a value of 0.4 (0.05*8.0) not less than 0.25 will be determined. Compare it with an additional alternative example with a larger difference (0.5) but the same duration of 2.0 seconds. In such an alternative example, a value of 0.3 (0.5*0.6) not less than 0.25 will be determined. Therefore, by considering duration, magnitude and / or other considerations, the occurrence of incorrectly determined false negatives can be mitigated. For example, considering duration can ensure that subsequent utterances do mean another attempt of previous utterances. In addition, for example, considering the magnitude can ensure that previous utterances may indeed be hot words, rather than another non-hot word utterance that occurs only before subsequent utterances. As another example, previous utterances and subsequent utterances may come from the same person to ensure that subsequent utterances do mean another attempt of previous utterances. Determining that utterances may come from the same person can be based on speaker identification technology and / or based on comparing the sound characteristics (e.g., pitch, intonation, and rhythm) of the two utterances.
[0018] As another example, assume again that the hotword detection model is trained to generate a prediction output indicating the probability of whether the hotword "good assistant" is present in the audio data, and assume that if the probability is greater than 0.85, the hotword will be determined to be present. Also assume again that the initial spoken utterance captured in the audio data includes the hotword "good assistant", but the prediction output generated based on processing the audio data indicates a probability of only 0.8. Further assume that another user interface input is received 1.5 seconds after the initial spoken utterance, and the other user interface input is an alternative call to the automated assistant, such as actuation of an explicit automated assistant call button (e.g., a hardware button or a software button), a sensed "squeeze" of the device (e.g., calling the automated assistant when squeezing the device with at least a threshold amount of force) or other explicit automated assistant calls. Therefore, in this example, the user initially says "good assistant" to call the assistant, which is not recognized as a call hotword, and the user quickly follows the call to the assistant in a replacement manner. In this example, based on 0.8 being less than 0.85, an initial decision will be made not to initiate certain processing of the audio data in response to the initial spoken utterance. However, certain processing of the audio data will be initiated in response to subsequent alternative calls. Additionally, in this example, it may be determined that the initial decision was incorrect. This may be based on another user interface input that actually invoked the assistant (i.e., a subsequent alternative invocation), may be based on the probability of the initial spoken utterance (0.8), and / or the duration (2.0 seconds) between receiving the initial spoken utterance and a subsequent another user interface input that contradicted the first spoken utterance by satisfying a threshold.
[0019] Although the above provides an example of a hot word detection model that monitors the presence of a "call" hot word, which will result in certain processing of audio data that follows within a threshold amount of time of the call hot word, it should be understood that the technology disclosed herein can be additionally or alternatively applied to other hot word detection models, which can be used to monitor words (including multi-word phrases) that, if determined to be present, will directly result in the execution of corresponding actions, at least under certain conditions.
[0020] For example, it is assumed that a hot word detection model trained to generate a predicted output is provided, the predicted output indicating the probability of whether a specific hot word such as "stop" and / or "suspend" exists in the audio data, and if the probability is greater than 0.85, the hot word is determined to exist. It is further assumed that the hot word detection model is active under specific conditions such as alarm sounding and / or music playing, and it is assumed that if the predicted output indicates that the hot word exists, any currently rendered automated assistant function that stops the audio output will be initiated. In other words, such a hot word detection model enables the user to simply say "stop" to cause the sounding alarm and / or the playing music to be suspended. It is also assumed that the initial spoken utterance captured in the audio data includes the hot word "stop", but the predicted output generated based on the processing of the audio data indicates a probability of only 0.8. It is further assumed that the subsequent spoken utterance captured in the additional audio data is received 0.5 seconds after the initial spoken utterance includes the hot word "stop" (e.g., after the initial spoken utterance is completed), and the predicted output generated based on the processing of the additional audio data indicates a probability of 0.9. In this example, it can be determined that the initial decision is incorrect. This can be based on another user interface input (i.e., a subsequent spoken utterance) satisfying the threshold, can be based on the probability of the initial spoken utterance (0.8), the duration between receiving the initial spoken utterance and the subsequent spoken utterance (which contradicts the first spoken utterance by satisfying the threshold) (0.5 seconds), and / or can be based on determining that the initial spoken utterance and the subsequent spoken utterance are likely from the same user (e.g., using a speaker identifier).
[0021] Some specific examples of determining whether a decision about whether to initiate a currently dormant automated assistant function is incorrect are now provided with reference to a continued conversation model, which is utilized to generate a prediction output based on processing audio data that captures a subsequent utterance, which prediction output can optionally be processed using a machine learning model together with processing recognition text from the subsequent utterance and / or its representation. For example, a first branch of the continued conversation model can be utilized to process audio data and generate a first branch output, a second branch of the continued conversation model can be utilized to process recognition text and / or its representation and generate a second branch output, and the prediction output can be based on processing both the first branch output and the second branch output. The prediction output can specify whether to initiate certain processing for a subsequent utterance, its recognition text, and / or its representation. For example, it can specify whether to attempt to generate an action based on recognition text and / or its representation and / or whether to perform an action.
[0022] As an example, assume that the continued conversation model is trained to generate a prediction output that indicates the probability that the subsequent utterance is intended for the automated assistant, and if the probability is greater than 0.80, the subsequent utterance will be determined to be intended for the automated assistant. It is also assumed that the initial subsequent utterance captured in the audio data includes "Remind me to take out the trash tomorrow", but the prediction output generated based on the processing using the continued conversation model indicates a probability of only 0.7. As a result, some processing based on the initial subsequent utterance is not performed. For example, a reminder of "tomorrow" for "taking out the trash" will not be generated. It is further assumed that the user then calls the assistant (e.g., using a hot word or using the assistant button) 2.5 seconds later and then provides a subsequent utterance of "remind me to take out the trash tomorrow". Since the subsequent utterance is provided after the call, it can be fully processed so that a reminder of "tomorrow" for "taking out the trash" is generated. Therefore, in this example, the user initially provides a subsequent utterance intended for the assistant, which is not identified as intended for the assistant, and the user quickly follows by calling the assistant (i.e., not in follow mode) and providing another instance of the utterance so that the utterance is fully processed by the assistant. In this example, it can be determined that the initial determination that the subsequent utterance was not intended for the assistant was incorrect. This can be based on one or more measures of similarity between the subsequent utterance and another user interface input (i.e., the subsequent spoken utterance), can be based on the probability of the initial subsequent utterance (0.7), and / or can be based on the duration between receiving the initial subsequent utterance and the call to provide the subsequent spoken utterance (2.5 seconds).
[0023] The measure of similarity may include, for example, duration similarity based on a comparison of the duration of the initial subsequent utterance and the duration of the subsequent spoken utterance, voice similarity based on a comparison of the voice characteristics of the initial subsequent utterance and the voice characteristics of the subsequent spoken utterance, and / or text similarity based on a comparison of the recognition text of the initial subsequent utterance and the recognition text of the subsequent spoken utterance. Generally, the greater the similarity, the greater the probability that the subsequent spoken utterance is determined to be the correction of the decision. For example, in this example, the initial subsequent utterance and the subsequent spoken utterance will have a high degree of duration similarity, voice similarity, and text similarity. Determining that the previous decision is incorrect can additionally or alternatively depend on the duration between receiving the initial subsequent utterance and the call being less than a threshold duration (e.g., 4.0 seconds or other threshold duration) and / or depending on the probability of the initial decision being within the range of a 0.80 threshold (e.g., within the range of 0.35 or other thresholds). In other words, in this case, an incorrect decision is determined only when the duration is less than the threshold duration and the probability is within the threshold range of the 0.80 threshold.
[0024] More generally, incorrect decisions may be determined based on duration and / or probability, without necessarily requiring that the decisions meet any corresponding threshold. As a non-limiting example, whether a decision is correct may be based on multiplying the difference between the probability for an initial decision and a threshold by: (1) a similarity metric (where the similarity metric is between 0 and 1, and larger values indicate greater similarity); and / or (2) a duration-based factor (where larger factor values correspond to larger durations), and determining whether the resulting value is less than a threshold.
[0025] Considering one or more of these factors can mitigate the occurrence of incorrectly determining false negatives and / or false positives. For example, considering a similarity metric will prevent determining a false negative where the initial follow-up utterance is “Remind me to take out the trash tomorrow” (and is intended for another person near the user, not for the automated assistant), and the follow-up utterance received after the subsequent call is “What is the square root of 256”. Also, for example, considering a similarity metric will prevent determining a false negative where the initial follow-up utterance is “Remind me to take out the trash tomorrow” (and is intended for another person near the user, not for the automated assistant), and the follow-up utterance received after the subsequent call is “Remind me to take out the trash tomorrow”, but the follow-up utterance is received 2 minutes after the initial follow-up utterance (e.g., only after the user later determines that this may be a good utterance to direct to the automated assistant).
[0026] Reference is now made to the invocation model without hot words to provide some specific examples of determining whether a decision about whether to initiate a currently dormant automated assistant function is incorrect. Under at least some conditions, the invocation model without hot words can be used to process data from one or more non-microphone sensors (and / or an abstraction of processing such data) to generate a prediction output that will initiate a currently dormant automated assistant function when it meets a threshold. For example, the invocation model without hot words can process visual data from a visual sensor of an automated assistant client device, and generate a prediction output that should meet a threshold in response to visual data including certain gestures of the user and / or in response to visual data including the gaze of the user directed at the automated assistant client device ("directional gaze") . For example, the invocation model without hot words can be used to invoke an automated assistant (e.g., in place of a hot word) in response to certain gestures (e.g., waving and / or thumbs up) and / or in response to a directional gaze of at least a threshold duration.
[0027] More generally, in various implementations, the invocation model without hot words can be utilized to monitor the presence of speechless physical movement (e.g., gestures or postures, eye gaze, facial movement or expression, mouth movement, proximity of the user to the client device, body posture or posture, and / or other speechless techniques) detected via one or more non-microphone sensor components of the client device. When detected, such speechless physical movement will initiate certain processing of sensor data that follows within a threshold amount of time of the speechless physical movement. One or more non-microphone sensors may include cameras or other visual sensors, proximity sensors, pressure sensors, accelerometers, magnetometers, and / or other sensors, and may be used to generate sensor data in addition to or in lieu of audio data captured via a microphone of the client device. In some implementations, speechless physical movement detected via one or more non-microphone sensor components of the client device may act as a proxy for explicit invocation hot words, and the user does not need to provide explicit verbal utterances including explicit invocation hot words. In other implementations, the speechless physical movement detected via one or more non-microphone sensor components of the client device is in addition to spoken speech captured via one or more microphones of the client device and including a hotword (such as “Hey assistant,” “Ok assistant,” “Assistant,” or any other suitable hotword).
[0028] When an automated assistant function is activated based on a predicted output generated using a call model without hot words, subsequent verbal utterances can be received and processed by the automated assistant, which will directly cause the corresponding actions included in the subsequent verbal utterances to be executed. Moreover, when the automated assistant function is activated, various human-perceivable prompts can be provided to indicate that the automated assistant function is activated. These human-perceivable prompts can include audible "bells", audible "verbal output" (e.g., "It looks like you are talking to the assistant"), visual symbols on the display screen of the client device, illumination of the light-emitting diodes of the client device, and / or other human-perceivable prompts indicating that the automated assistant function is activated.
[0029] As a specific example, assume that the invocation model without hot words is trained to generate a prediction output indicating the probability of whether a speechless physical movement is detected in the sensor data, and assume that if the probability is greater than 0.85, the speechless physical movement will be determined to be used as a proxy for the hot word. It is also assumed that the prediction output generated based on the sensor data indicates a probability of only 0.80 based on detecting the following: (1) movement of the user's mouth (also referred to as "mouth movement" in this article); and / or (2) directing the user's gaze at the client device (also referred to as "directional gaze" in this article). It is further assumed that a spoken utterance captured in the audio data via the microphone of the client device is received 2.0 seconds after the mouth movement and / or directional gaze, and the spoken utterance includes the invocation hot word "good assistant", and the prediction output generated based on processing the additional audio data indicates a probability of 0.90. Therefore, in this example, the mouth movement and / or directional gaze initially used by the user to attempt to invoke the assistant is not identified as acting as a proxy for the invocation hot word, and the user quickly follows the spoken utterance of "good assistant" to attempt to invoke the assistant again—and this follow-up instance is identified as including the invocation hot word. In this example, based on 0.80 being less than 0.85, an initial decision will be made not to initiate certain processing of the audio data in response to mouth movement and / or directional gaze. However, based on 0.90 being greater than 0.85, a subsequent decision will be made to initiate certain processing of the audio data in response to spoken utterances. In addition, in this example, it can be determined that the initial decision is incorrect (i.e., the decision is a false negative). This can be based on another user interface input (i.e., spoken utterance) that meets the threshold, can be based on the probability of initial mouth movement and / or directional gaze (e.g., 0.80), and / or the duration (2.0 seconds) between receiving initial mouth movement and / or directional gaze and spoken utterances (contradicting the initial decision that meets the threshold), as described herein. For example, determining that the previous decision is incorrect can be a function of the probability that the duration is less than a threshold duration (e.g., 4.0 seconds or other threshold duration) and the initial decision is within the range of a 0.85 threshold (e.g., within a range of 0.35 or other ranges). In other words, in this case, an incorrect decision is determined only when the duration is less than the threshold duration and the probability is within the threshold range of the 0.85 threshold.
[0030] In contrast, assume again that the invocation model without hot words is trained to generate a prediction output indicating the probability of whether a speechless physical movement is detected in the sensor data, and assume that if the probability is greater than 0.85, the speechless physical movement will be determined to act as a proxy for the hot word. Assume again that the prediction output generated based on the sensor data indicates a probability of 0.90 based on the detection of the following: (1) mouth movement; and / or (2) directional gaze. Therefore, based on 0.90 being greater than 0.85, an initial decision will be made to initiate certain automated assistant functions for processing subsequent spoken utterances following the mouth movement and / or directional gaze, and the client device can provide a given human-perceivable prompt that the client device has initiated certain automated assistant functions for processing subsequent spoken utterances following the mouth movement and / or directional gaze. Further assume that another user interface input is received after the mouth movement and / or directional gaze, and the other user interface input contradicts the initial decision. Therefore, in this example, the user initially directs the mouth movement and / or directional gaze to the client device, which is identified as acting as a proxy for invoking the hot word, and the user quickly follows up with another user interface input that cancels the invocation of the assistant. In this example, based on 0.90 being greater than 0.85, an initial decision will be made to initiate certain automated assistant functions for processing subsequent audio data that follows mouth movements and / or directional gazes. However, based on another user interface input, a subsequent decision will be made to deactivate or turn off certain automated assistant functions initiated for processing subsequent audio data that follows mouth movements and / or directional gazes. In addition, in this example, it can be determined that the initial decision is incorrect (i.e., the decision is a false positive) based on another user interface input that contradicts the initial decision. The other user interface input can be a physical movement without additional words that contradicts the initial decision, a verbal utterance (e.g., "stop", "no", and / or other verbal utterances that contradict the initial decision), an explicit automated assistant invocation button (e.g., a hardware button or a software button) that negates the initial decision, a sensed "squeeze" of a device that negates the initial decision (e.g., when the device is squeezed with at least a threshold amount of force, the invocation of the automated assistant is negated), and / or another other user interface input that contradicts the initial decision. Moreover, determining that the initial decision is incorrect based on another user interface input that contradicts the initial decision can be based on another predicted output of another user interface input that fails to meet a threshold determined using the hot-word-free invocation model disclosed herein or another machine learning model, can be based on a duration between receiving the initial mouth movement and / or directional gaze and the other user interface input as described herein, and / or can be based on a probability for the initial decision being within a range of a 0.85 threshold (e.g., within a range of 0.35 or other ranges), as described herein.
[0031] As another example, assume again that the invocation model without hot words is trained to generate a prediction output that indicates a probability of whether a speechless physical movement is detected in the sensor data, and if the probability is greater than 0.85, the speechless physical movement will be determined to act as a proxy for the hot word. Assume again that the prediction output generated based on the sensor data indicates a probability of only 0.80 based on detecting: (1) the user's proximity to the client device (e.g., the user is within a threshold distance of the client device); and / or (2) a gesture directing the user at the client device (e.g., as indicated by hand movements or gestures, body language or gestures, and / or other gesture indications). Assume further that when the user is within the threshold distance of the client device, another user interface input is received 1.5 seconds after the user's initial gesture, and the other user interface input is an alternative invocation of the automated assistant, such as actuation of an explicit automated assistant invocation button (e.g., a hardware button or a software button), a sensed "squeeze" of the device (e.g., when the device is squeezed with at least a threshold amount of force, it invokes the automated assistant), or other explicit automated assistant invocation. Therefore, in this example, when the user is within the threshold distance of the client device to call the assistant, the user initially provides a gesture that is not recognized as acting as a proxy for calling a hot word, and the user quickly follows the call assistant in an alternative manner. In this example, based on 0.80 being less than 0.85, an initial decision will be made in response to the initial spoken utterance without initiating certain processing of the audio data. However, in response to the alternative call, certain processing of the audio data will be initiated. In addition, in this example, it can be determined that the initial decision is incorrect (i.e., the decision is a false negative). This can be based on another user interface input (i.e., an alternative call) that actually calls the assistant, can be based on the probability (e.g., 0.80) of the user's initial gesture when the user is within the threshold distance of the client device, and / or can be based on the duration (e.g., 2.0 seconds) between receiving the user's initial gesture and another subsequent user interface input (contradicting the first gesture by meeting the threshold) as described herein when the user is within the threshold distance of the client device.
[0032] Although the above provides an example of an invocation model without a hot word that monitors the presence of speechless physical movement that acts as a proxy for an "invoke" hot word, which when detected will result in subsequent spoken utterances following the speechless physical movement within a threshold amount of time and / or certain processing of the physical movement without subsequent utterances, it should be understood that the techniques disclosed herein may additionally or alternatively be applied to speechless physical movement provided in conjunction with spoken utterances that include the "invoke" hot word, and may be used to monitor speechless physical movement and / or the presence of words (including multi-word phrases) under at least certain conditions, which, if determined to be present, will directly result in the execution of a corresponding action.
[0033] Moreover, it should be understood that the technology disclosed herein can be additionally or alternatively applied to speechless physical movements including actions performed in response thereto, such as gestures for "stop" for a client device when an alarm is sounding or music is playing, lifting up or down to control the volume of the client device, and other gestures for controlling the client device without providing any verbal words. For example, it is assumed that a call model without hot words is provided that is trained to generate a prediction output, which indicates the probability of whether there are certain speechless physical movements such as hand movements and / or gestures corresponding to "stop" and / or "abort" in the sensor data, and if the probability is greater than 0.85, it will be determined that there are certain speechless physical movements. It is further assumed that the call model without hot words is active under certain conditions such as alarm sounding and / or music playing, and it is assumed that if the prediction output indicates that there is a speechless movement, any currently rendered automated assistant function that stops the audio output will be initiated. In other words, such a call model without hot words enables the user to simply provide a hand movement and / or gesture corresponding to "stop" to cause the sounding alarm and / or playing music to be aborted.
[0034] In some implementations, when making a decision about whether to initiate a currently dormant automated assistant function and / or whether to turn off a currently active automated assistant function, a given client device may transmit audio data and / or other sensor data to a cloud-based machine learning model (e.g., a cloud-based hot word detection engine, a cloud-based continued conversation engine, a cloud-based hot word-free call engine, and / or other cloud-based engines). Cloud-based machine learning models are typically more robust than machine learning models on devices and can be used to verify decisions made at a given client device. In some versions of those implementations, a cloud-based machine learning model can process audio data and / or other sensor data, can make a determination about whether a client device should initiate a specific automated assistant function, and can transmit to a given client device an indication of whether the decision made at a given client device is correct or incorrect (i.e., whether the decision is a false negative or a false positive). In some other versions of those implementations, a given client device may initiate a currently dormant automated assistant function and / or turn off a currently active automated assistant function based on an indication received from a cloud-based machine learning model. For example, if a given client device makes a decision not to initiate a currently dormant automated assistant function based on a predicted output generated using an on-device machine learning model, transmits audio data and / or sensor data used to generate the predicted output to a cloud-based machine learning model, and receives an indication from the cloud-based machine learning model that the decision is incorrect (i.e., the decision is a false negative), then the client device can initiate the currently dormant automated assistant function and use this instance to generate gradients for training the on-device machine learning model. As another example, if a given client device makes a decision to initiate a currently dormant automated assistant function based on a predicted output generated using an on-device machine learning model, transmits audio data and / or sensor data used to generate the predicted output to a cloud-based machine learning model, and receives an indication from the cloud-based machine learning model that the decision is incorrect (i.e., the decision is a false positive), then the client device can turn off the currently active automated assistant function and use this instance to generate gradients for training the on-device machine learning model. Therefore, in these implementations, in addition to the on-device machine learning model, the cloud-based machine learning model can be used to identify false negatives and / or false positives.
[0035] By utilizing one or more techniques described herein, based on audio data corresponding to spoken speech and / or sensor data corresponding to physical movement without speech, the occurrence of false negatives and / or false positives can be automatically identified and marked locally at the corresponding client device. In addition, the identified and marked false positives and false negatives can be used locally at the corresponding client device to generate gradients. The gradient can be used locally at the corresponding client device to update the corresponding locally stored machine learning model and / or can be transmitted to the remote system for use in updating the corresponding global model. This results in improvements in the performance of the corresponding locally stored machine learning model and / or the corresponding global model (which can be transmitted to various client devices for use).
[0036] Additionally or alternatively, automatically marking false positives and / or false negatives locally at the corresponding client device can maintain the privacy of user data (e.g., spoken utterances, etc.) because these user data may never be transmitted from the corresponding client device and / or will be marked without any manual review. Moreover, such automatic marking can save various resources, such as network resources that would otherwise be needed to transmit corresponding data (e.g., bandwidth-intensive audio data and / or visual data) to the client device of a human reviewer for marking and / or client device resources that would otherwise be used to review and manually mark the corresponding data. In addition, with current manual marking technology, the occurrence of false negatives may never be transmitted from the client device to the server for manual review and marking. Therefore, with current technology, it may never be possible to train (or only minimally train) a machine learning model based on the actual real-world occurrence of false negatives. However, the implementation disclosed herein enables automatic identification and marking of false negatives at the client device, generates gradients based on such false negatives, and updates the corresponding machine learning model based on the generated gradients.
[0037] The above description is provided as an overview of some implementations of the present disclosure. Further descriptions of those implementations and other implementations are described in more detail below.
[0038] Various implementations may include a non-transitory computer-readable storage medium storing instructions that are executed by one or more processors (e.g., a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), and / or a tensor processing unit (TPU)) to perform a method such as one or more of the methods described herein. Other implementations may include an automated assistant client device (e.g., a client device that includes at least an automated assistant interface for interfacing with a cloud-based automated assistant component), the automated assistant client device including a processor that is operable to execute the stored instructions to perform a method such as one or more of the methods described herein. Additional implementations may also include a system of one or more servers including one or more processors that are operable to execute the stored instructions to perform a method such as one or more of the methods described herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1A , 1B , 1C and 1D depict example process flows illustrating various aspects of the present disclosure according to various implementations.
[0040] Figure 2 Depicted include Figures 1A-1D 00106] A block diagram of various components of the EMBODIMENTS 100 and an example environment in which implementations disclosed herein may be implemented.
[0041] Figure 3 A flow chart illustrating an example method of generating gradients locally on a client device based on false negatives and transmitting the gradients to a remote server and / or utilizing the generated gradients to update weights of a speech recognition model on the device is depicted.
[0042] Figure 4 A flow diagram illustrating an example method of generating gradients locally on a client device based on false positives and transmitting the gradients to a remote server and / or utilizing the generated gradients to update weights of a speech recognition model on the device is depicted.
[0043] Figure 5 A flow chart illustrating an example method of updating weights of a global speech recognition model based on gradients received from a remote client device and transmitting the updated weights or the updated global speech recognition model to the remote client device is depicted.
[0044] Figure 6 An example architecture for a computing device is depicted. DETAILED DESCRIPTION
[0045] Figures 1A-1D An example process flow illustrating various aspects of the present disclosure is depicted. The client device 110 Figure 1A is illustrated in and includes a Figure 1A Components within the box. The machine learning engine 122A can receive audio data 101 corresponding to spoken words detected via one or more microphones of the client device 110 and / or other sensor data 102 corresponding to physical movements (e.g., gestures and / or movements, body gestures and / or body movements, eye gaze, facial movements, mouth movements, etc.) without words detected via one or more non-microphone sensor components of the client device 110. One or more non-microphone sensors may include a camera or other visual sensor, a proximity sensor, a pressure sensor, an accelerometer, a magnetometer, and / or other sensors. The machine learning engine 122A uses the machine learning model 152A to process the audio data 101 and / or other sensor data 102 to generate a predicted output 103. As described herein, the machine learning engine 122A can be a hot word detection engine 122B, a call engine 122C without hot words, a continued conversation engine 122D, and an alternative engine, such as a voice activity detector (VAD) engine, an endpoint detector engine, and / or other engines.
[0046] In some implementations, when the machine learning engine 122A generates the prediction output 103, it may be stored locally on the client device in the on-device storage 111, and optionally associated with the corresponding audio data 101 and / or other sensor data 102. In some versions of those implementations, the prediction output may be retrieved by the gradient engine 126 for use in generating the gradient 106 at a later time, such as when one or more conditions described herein are met. The on-device storage 111 may include, for example, a read-only memory (ROM) and / or a random access memory (RAM). In other implementations, the prediction output 103 may be provided to the gradient engine 126 in real time.
[0047] Client device 110 may make a decision based on determining at block 182 whether predicted output 103 satisfies a threshold whether to initiate a currently dormant automated assistant function (e.g., Figure 2Automated assistant 295), avoid initiating a currently dormant automated assistant function, and / or use assistant activation engine 124 to shut down a currently active automated assistant function. Automated assistant functions may include: speech recognition that generates recognized text, generating natural language understanding (NLU) output of NLU, generating a response based on the recognized text and / or NLU output, transmitting audio data to a remote server, and / or transmitting recognized text to a remote server. For example, assuming that the predicted output 103 is a probability (e.g., 0.80 or 0.90) and the threshold at box 182 is a threshold probability (e.g., 0.85), if the client device 110 determines at box 182 that the predicted output 103 (e.g., 0.90) meets the threshold (e.g., 0.85), the assistant activation engine 124 can initiate the currently dormant automated assistant function.
[0048] In some implementations, and as Figure 1B , machine learning engine 122A may be hot word detection engine 122B. It is noteworthy that various automated assistant functions such as on-device speech recognizer 142, on-device NLU engine 144, and / or on-device fulfillment engine 146 are currently dormant (i.e., as indicated by dashed lines). In addition, assume that the predicted output 103 generated using hot word detection model 152B and based on audio data 101 satisfies the threshold at box 182, and the voice activity detector 128 detects user speech directed to client device 110.
[0049] In some versions of these implementations, assistant activation engine 124 activates on-device speech recognizer 142, on-device NLU engine 144, and / or on-device fulfillment engine 146 as currently dormant automated assistant functions. For example, on-device speech recognizer 142 can use on-device speech recognition model 142A to process audio data 101 for spoken utterance to generate recognition text 143A, the spoken utterance including the hot word "good assistant" and additional commands and / or phrases following the hot word "good assistant", on-device NLU engine 144 can use on-device NLU model 144A to process recognition text 143A to generate NLU data 145A, on-device fulfillment engine 146 can use on-device fulfillment model 146A to process NLU data 145A to generate fulfillment data 147A, and client device 110 can use fulfillment data 147A when performing 150 one or more actions in response to audio data 101.
[0050] In other versions of these implementations, assistant activation engine 124 activates only on-device fulfillment engine 146, without activating on-device speech recognizer 142 and on-device NLU engine 144, to process various commands, such as “no,” “stop,” “cancel,” and / or other commands that can be processed without on-device speech recognizer 142 and on-device NLU engine 144. For example, on-device fulfillment engine 146 processes audio data 101 using on-device fulfillment model 146A to generate fulfillment data 147A, and client device 110 can use fulfillment data 147A in performing 150 one or more actions in response to audio data 101. Furthermore, in versions of these implementations, the assistant activation engine 124 may initially activate a currently dormant automation functionality to verify that the decision made at box 182 is correct (e.g., the audio data 101 actually includes the hot word “good assistant”) by initially activating only the on-device speech recognizer 142 to determine that the audio data 101 includes the hot word “good assistant”, and / or the assistant activation engine 124 may transmit the audio data 101 to one or more servers (e.g., remote system 160) to verify that the decision made at box 182 is correct (e.g., the audio data 101 actually includes the hot word “good assistant”).
[0051] In some implementations, and as Figure 1C As depicted in , machine learning engine 122A may be a hot-word-free invocation engine 122C. It is noteworthy that various automated assistant functions such as the on-device speech recognizer 142, the on-device NLU engine 144, and / or the on-device fulfillment engine 146 are currently dormant (i.e., as indicated by the dashed lines). In addition, it is assumed that the predicted output 103 generated using the hot-word-free invocation model 152C and based on other sensor data 102 satisfies the threshold at box 182, and the voice activity detector 128 detects user speech directed to the client device 110.
[0052] In some versions of these implementations, assistant activation engine 124 activates on-device speech recognizer 142, on-device NLU engine 144, and / or on-device fulfillment engine 146 as currently dormant automated assistant functions. For example, in response to activating an automated assistant function for speechless physical movement acting as an agent for a hot word, on-device speech recognizer 142 can use on-device speech recognition model 142A to process commands and / or phrases that occur with and / or follow the speechless physical movement that acts as an agent for a hot word to generate recognition text 143A, on-device NLU engine 144 can use on-device NLU model 144A to process recognition text 143A to generate NLU data 145A, on-device fulfillment engine 146 can use on-device fulfillment model 146A to process NLU data 145A to generate fulfillment data 147A, and client device 110 can use fulfillment data 147A when performing 150 one or more actions in response to audio data 101.
[0053] In other versions of these implementations, assistant activation engine 124 activates only on-device fulfillment engine 146, without activating on-device speech recognizer 142 and on-device NLU engine 144, to process various commands, such as "no," "stop," "cancel," and / or other commands. For example, on-device fulfillment engine 146 processes the command or phrase that occurs with and / or follows the speechless physical movement using on-device fulfillment model 146A to generate fulfillment data 147A, and client device 110 may use fulfillment data 147A when performing 150 one or more actions in response to the command or phrase that occurs with and / or follows the speechless physical movement.
[0054] Furthermore, in some versions of these implementations, assistant activation engine 124 may initially activate currently dormant automation functionality by initially only activating speech recognizer 142 on the device to determine that the command and / or phrase occurring with and / or following the speechless physical movement is intended for the assistant to verify that the determination made at box 182 is correct (e.g., that the speechless physical movement captured by other sensor data 102 is in fact intended to act as a proxy for the hotword), and / or assistant activation engine 124 may transmit other sensor data 102 to one or more servers (e.g., remote system 160) to verify that the determination made at box 182 is correct (e.g., that the speechless physical movement captured by other sensor data 102 is in fact intended to act as a proxy for the hotword).
[0055] In some implementations, and as Figure 1D, machine learning engine 122A is a continued conversation engine 122D. Notably, various automated assistant functions such as an on-device speech recognizer 142 and an on-device NLU engine 144 are already active (i.e., as indicated by solid lines) as a result of previous interactions with the assistant (e.g., as a result of an utterance that includes a hot word and / or a speechless physical movement that acts as an agent for the hot word). Client device 110 may retrieve (e.g., from storage 111 on the device) recognized text 143A and / or NLU data 145A from these previous interactions. For example, additional audio data 101A capturing a subsequent spoken utterance may be received (i.e., after a spoken utterance that includes a hot word and / or a speechless physical movement that acts as an agent for a hot word that triggers the automated assistant), and may be a follow-up request, a clarifying response, a response to a prompt from the automated assistant, an additional user request, and / or other interaction with the automated assistant. Furthermore, assume that the predicted output 103 generated using the continued conversation model 152D and based on the additional audio data 101A, the recognized text 143A from the previous interaction, and the NLU data 145A from the previous interaction satisfies the threshold at box 182 and the voice activity detector 128 detects user speech directed to the client device 110.
[0056] In some versions of these implementations, assistant activation engine 124 suppresses shutting down on-device speech recognizer 142 and on-device NLU engine 144, and activates on-device fulfillment engine 146 as a currently dormant automated assistant function. For example, on-device speech recognizer 142 may use on-device speech recognition model 142A to process additional audio data 101A for subsequent spoken utterances omitting the hot word "good assistant" to generate another recognition text 143B, on-device NLU engine 144 may use on-device NLU model 144A to process another recognition text 143B to generate another NLU data 145B, on-device fulfillment engine 146 may use on-device fulfillment model 146A to process another NLU data 145B to generate another fulfillment data 147B, and client device 110 may use another fulfillment data 147A when performing 150 one or more actions in response to additional audio data 101A.
[0057] Moreover, in some versions of these implementations, assistant activation engine 124 may initially activate currently dormant automated assistant functionality by initially processing additional audio data 101A using only speech recognizer 142 on the device to verify that the decision made at box 182 is correct (e.g., additional audio data 101A is in fact intended for an automated assistant), and / or assistant activation engine 124 may transmit additional audio data 101A to one or more servers (e.g., remote system 160) to verify that the decision made at box 182 is correct (e.g., additional audio data 101A is in fact intended for an automated assistant).
[0058] Return to Figure 1A If client device 110 determines at block 182 that prediction output 103 (e.g., 0.80) fails to satisfy a threshold value (e.g., 0.85), assistant activation engine 124 may suppress initiation of currently dormant automated assistant functions and / or shut down any currently active automated assistant functions. Additionally, if client device 110 determines at block 182 that prediction output 103 (e.g., 0.80) fails to satisfy a threshold value (e.g., 0.85), client device 110 may determine at block 184 whether another user interface input is received. For example, the another user interface input may be an additional spoken utterance including a hot word, a physical movement without additional utterances that acts as a proxy for the hot word, an actuation of an explicit automated assistant invocation button (e.g., a hardware button or a software button), a sensed “squeeze” of a client device 110 device (e.g., invoking an automated assistant when the client device 110 is squeezed with at least a threshold amount of force), and / or other explicit automated assistant invocations. If the client device 110 determines at block 184 that another user interface input has not been received, the client device 110 may stop identifying the correction and end at block 190 .
[0059] However, if the client device 110 determines at block 184 that another user interface input was received, the system may determine whether the other user interface input received at block 184 contradicts the decision made at block 182 (including the correction at block 186). The correction may identify a false negative (e.g., Figure 3 more fully described) or false positives (e.g., as discussed in Figure 4 184 does not include a correction, the client device 110 may stop identifying the correction and end at block 190. However, if the client device 110 determines at block 186 that the other user interface input received at block 184 includes a correction that contradicts the initial decision made at block 182, the client device 110 may determine the ground truth output 105.
[0060] As a non-limiting example of a false negative, assume that machine learning engine 122A is trained to generate a probability as prediction output 103, assume that client device 110 incorrectly determines at box 182 that prediction output 103 (e.g., 0.80) fails to meet a threshold (e.g., 0.85), and assume that client device 110 refrains from initiating a currently dormant automated assistant function and / or refrains from closing a currently active automated assistant function. Further, assume that client device 110 determines based on another user interface input received at box 184 that another user interface input contradicts the initial decision made at box 182, and that client device 110 should have initiated a currently dormant automated assistant function and / or refrained from closing a currently active automated assistant function. In this case, ground truth output 105 may also be a probability (e.g., 1.00) indicating that client device 110 should have initiated a currently dormant automated assistant function and / or refrained from closing a currently active automated assistant function.
[0061] As a non-limiting example of a false positive, assume that machine learning engine 122A is trained to generate a probability as prediction output 103, client device 110 incorrectly determines at box 182 that prediction output 103 (e.g., 0.90) satisfies a threshold (e.g., 0.85), and client device 110 initiates a currently dormant automated assistant function and / or inhibits closing a currently active automated assistant function. Further, assume that client device 110 determines based on another user interface input received at box 184 that another user interface input contradicts the initial decision made at box 182, and client device 110 should not have initiated a currently dormant automated assistant function and / or inhibited closing a currently active automated assistant function. In this case, ground truth output 105 may also be a probability (e.g., 0.00) indicating that client device 110 should not have initiated a currently dormant automated assistant function and / or inhibited closing a currently active automated assistant function. Although the predicted output 103 and the ground truth output 105 are described herein as probabilities, it should be understood that this is not meant to be limiting and the predicted output 103 and the ground truth output 105 may be labels, annotations, binary values, and / or other likelihood measures.
[0062] In some implementations, the gradient engine 126 can generate the gradient 106 based on the predicted output 103 to the ground truth output 105. For example, the gradient engine 126 can generate the gradient 106 based on comparing the predicted output 103 to the ground truth output 105. In some versions of those implementations, the client device 110 stores the predicted output 103 and the corresponding ground truth output 105 locally in the storage 111 on the device, and the gradient engine 126 retrieves the predicted output 103 and the corresponding ground truth output 105 to generate the gradient 106 when one or more conditions are met. The one or more conditions may include, for example, that the client device is charging, the client device has at least a threshold charging state, the temperature of the client device (based on one or more temperature sensors on the device) is less than a threshold, and / or the client device is not held by the user. In other versions of those implementations, the client device 110 provides the predicted output 103 and the ground truth output 105 to the gradient engine 126 in real time, and the gradient engine 126 generates the gradient 106 in real time.
[0063] Moreover, the gradient engine 126 can provide the generated gradients 106 to the on-device machine learning training engine 132A. The on-device machine learning training engine 132A uses the gradients 106 when it receives the gradients 106 to update the on-device machine learning model 152A. For example, the on-device machine learning training engine 132A can utilize back propagation and / or other techniques to update the on-device machine learning model 152A. Note that in some implementations, the on-device machine learning training engine 132A can utilize batch processing techniques to update the on-device machine learning model 152A based on the gradients 106 and additional gradients determined locally on the client device 110 based on the additional corrections.
[0064] In addition, the client device 110 may transmit the generated gradients 106 to the remote system 160. When the remote system 160 receives the gradients 106, the remote training engine 162 of the remote system 160 updates the global weights of the global speech recognition model 152A1 using the gradients 106 and the additional gradients 107 from the additional client devices 170. The additional gradients 107 from the additional client devices 170 may each be generated based on the same or similar techniques as described above with respect to the gradients 106 (but based on corrections to local identifications specific to those client devices).
[0065] The update distribution engine 164 can provide updated global weights and / or the updated global speech recognition model itself to the client device 110 and / or other client devices in response to satisfying one or more conditions, as indicated by 108. One or more conditions may include, for example, a threshold duration and / or amount of training since the last provision of updated weights and / or updated speech recognition models. One or more conditions may additionally or alternatively include, for example, an improvement in a metric of the updated speech recognition model and / or the passage of a threshold duration since the last provision of updated weights and / or updated speech recognition models. When the updated weights are provided to the client device 110, the client device 110 can replace the weights of the machine learning model 152A on the device with the updated weights. When the updated global speech recognition model is provided to the client device 110, the client device 110 can replace the machine learning model 152A on the device with the updated global speech recognition model.
[0066] In some implementations, the on-device machine learning model 152A is transmitted (e.g., via the remote system 160 or other component) for storage and use at the client device 110 based on the geographic region and / or other attributes of the client device 110 and / or the user of the client device 110. For example, the on-device machine learning model 152A may be one of N available machine learning models for a given language, but may be trained based on corrections that are specific to a particular geographic region, and provided to the client device 110 based on the client device 110 being primarily located in the particular geographic region.
[0067] Now turn to Figure 2 , in which Figures 1A-1D The client device 110 is illustrated in an implementation in which a machine learning engine on various devices is included as part of (or in communication with) the automated assistant client 240. Figures 1A-1D The corresponding machine learning models connected to the machine learning engines on various devices. For simplicity, Figure 2 Not shown Figures 1A-1D other parts of the device. Figure 2 Picture shows Figures 1A-1D 2 is an example of how machine learning engines on various devices and their corresponding machine learning models can be used by automated assistant client 240 to perform various actions.
[0068] Figure 2The client device 110 in FIG. 1 is illustrated with one or more microphones 211, one or more speakers 212, one or more cameras and / or other visual components 213, and a display 214 (e.g., a touch-sensitive display). The client device 110 may also include pressure sensors, proximity sensors, accelerometers, magnetometers, and / or other sensors for generating other sensor data in addition to the audio data captured by the one or more microphones 211. The client device 110 selectively executes at least an automated assistant client 240. Figure 2 In the example of , automated assistant client 240 includes hotword detection engine 122B on the device, invocation engine 122C without hotwords on the device, continued conversation engine 122D, speech recognizer 142 on the device, natural language understanding (NLU) engine 144 on the device, and fulfillment engine 146 on the device. Automated assistant client 240 also includes voice capture engine 242 and visual capture engine 244. Automated assistant client 240 may include additional and / or alternative engines, such as a voice activity detector (VAD) engine, an endpoint detector engine, and / or other engines.
[0069] One or more cloud-based automated assistant components 280 may optionally be implemented on one or more computing systems (collectively referred to as "cloud" computing systems) that are communicatively coupled to client device 110 via one or more local area networks and / or wide area networks (e.g., the Internet) generally indicated at 290. Cloud-based automated assistant components 280 may be implemented, for example, via a high-performance server cluster.
[0070] In various implementations, an instance of automated assistant client 240, through its interaction with one or more cloud-based automated assistant components 280, can form a logical instance of automated assistant 295 that appears to the user as being capable of human interaction (e.g., verbal interaction, gesture-based interaction, and / or touch-based interaction) with the user.
[0071] Client device 110 may be, for example, a desktop computing device, a laptop computing device, a tablet computing device, a mobile phone computing device, a computing device of a user's vehicle (e.g., an in-vehicle communication system, an in-vehicle entertainment system, an in-vehicle navigation system), a standalone interactive speaker, a smart appliance such as a smart TV (or a standard TV equipped with a networked dongle with automated assistant capabilities), and / or a wearable device of a user including a computing device (e.g., a watch of the user with the computing device, glasses of the user with the computing device, a virtual or augmented reality computing device). Additional and / or alternative client devices may be provided.
[0072] The one or more visual components 213 can take various forms, such as a single-view camera, a stereoscopic camera, a LIDAR component (or other laser-based component), a radar component, etc. The one or more visual components 213 can be used, for example, by the visual capture engine 242 to capture visual frames (e.g., image frames, laser-based visual frames) of the environment in which the client device 110 is deployed. In some implementations, such visual frames can be used to determine whether a user is present near the client device 110 and / or the distance of the user (e.g., the user's face) relative to the client device 110. Such determination can be used, for example, to determine whether to activate a user's facial recognition function. Figure 2 The machine learning engines and / or other engines on the various devices depicted in FIG.
[0073] The voice capture engine 242 may be configured to capture the user's voice and / or other audio data captured via the microphone 211. In addition, the client device 110 may include a pressure sensor, a proximity sensor, an accelerometer, a magnetometer, and / or other sensors for generating other sensor data in addition to the audio data captured via the microphone 211. As described herein, such audio data and other sensor data may be used by the hot word detection engine 122B, the call engine 122C without hot words, the continued conversation engine 122D, and / or other engines to determine whether to initiate one or more currently dormant automated assistant functions, suppress the initiation of one or more currently dormant automated assistant functions, and / or close one or more currently active automated assistant functions. The automated assistant function may include a speech recognizer 142 on the device, an NLU engine 144 on the device, a fulfillment engine 146 on the device, and additional and / or alternative engines. For example, the speech recognizer 142 on the device may utilize the speech recognition model 142A on the device to process the audio data capturing the spoken utterance to generate a recognition text 143A corresponding to the spoken utterance. The NLU engine 144 on the device optionally uses the NLU model 144A on the device to perform on-device natural language understanding on the recognized text 143A to generate NLU data 145A. The NLU data 145A may include, for example, an intent corresponding to the spoken utterance and optionally include parameters (e.g., slot values) for the intent. In addition, the fulfillment engine 146 on the device optionally uses the fulfillment model 146A on the device to generate fulfillment data 147A based on the NLU data 145A. The fulfillment data 147A may define local and / or remote responses (e.g., replies) to the spoken utterance, interactions performed with locally installed applications based on the spoken utterance, commands transmitted to an Internet of Things (IoT) device based on the spoken utterance (directly or via a corresponding remote system), and / or other resolution actions performed based on the spoken utterance. The fulfillment data 147A is then provided for local and / or remote execution / execution of the determined actions to parse the spoken utterance. Execution may include, for example, rendering local and / or remote responses (e.g., visually and / or audibly (optionally utilizing a local text-to-speech module)), interacting with locally installed applications, transmitting commands to IoT devices, and / or other actions.
[0074] Display 214 can be used for recognized text 143A and / or another recognized text 143B from speech recognizer 122 on the device, and / or one or more results from execution 150. Display 214 can also be one of the user interface output components through which the visual portion of the response from automated assistant client 240 is rendered.
[0075] In some implementations, the cloud-based automated assistant component 280 may include a remote ASR engine 282 that performs speech recognition, a remote NLU engine 282 that performs natural language understanding, and / or a remote fulfillment engine 284 that generates fulfillment. A remote execution module may also be optionally included, which implements remote execution based on fulfillment data determined locally or remotely. Additional and / or alternative remote engines may be included. As described herein, in various implementations, speech processing on the device, NLU on the device, fulfillment on the device, and / or execution on the device may be prioritized at least due to the delay and / or network usage reduction they provide when parsing spoken utterances (due to the fact that no client-server round trip is required to parse spoken utterances). However, one or more cloud-based automated assistant components 280 may be at least selectively utilized. For example, such components may be used in parallel with components on the device, and the output of such components may be used when local components fail. For example, the fulfillment engine 246 on the device may fail in some cases (e.g., due to the relatively limited resources of the client device 110), and in this case the remote fulfillment engine 283 may generate fulfillment data using the more robust resources of the cloud. Remote fulfillment engine 283 may operate in parallel with on-device fulfillment engine 246 and utilize its results when on-device fulfillment fails, or may be invoked in response to determining that on-device fulfillment engine 246 has failed.
[0076] In various implementations, an NLU engine (on-device and / or remote) may generate an NLU output that includes one or more annotations identifying text and one or more (e.g., all) terms of a natural language input. In some implementations, the NLU engine is configured to identify and annotate various types of grammatical information in the natural language input. For example, the NLU engine may include a morpheme module that may separate individual words into morphemes and / or annotate morphemes, for example, with their classes. The NLU engine may also include a portion of a speech tagger that is configured to annotate terms with their grammatical roles. Further, for example, in some implementations, the NLU engine may additionally and / or alternatively include a dependency parser configured to determine syntactic relationships between terms in the natural language input.
[0077] In some implementations, the NLU engine may additionally and / or alternatively include an entity tagger configured to annotate entity references in one or more snippets, such as references to people (including, for example, literary characters, celebrities, public figures, etc.), organizations, places (real and imagined), etc. In some implementations, the NLU engine may additionally and / or alternatively include a co-reference resolver (not depicted) configured to group or "cluster" references to the same entity based on one or more contextual cues. In some implementations, one or more components of the NLU engine may rely on annotations from one or more other components of the NLU engine.
[0078] The NLU engine may also include an intent matcher configured to determine the intent of a user engaging in an interaction with the automated assistant 295. The intent matcher may use various techniques to determine the intent of the user. In some implementations, the intent matcher may access one or more local and / or remote data structures that include, for example, multiple mappings between grammars and response intents. For example, the grammars included in the mappings may be selected and / or learned over time and may represent common intents of users. For example, a grammar "play <artist>"[play <artist>]" may be mapped to an intent that invokes a response action that causes music by <artist> to be played on client device 110. Another grammar "[weather|forecast]today" may be able to match user queries such as "what's the weather today" and "what's the forecast for today?" In addition to or in lieu of grammars, in some implementations, the intent matcher may use one or more trained machine learning models, alone or in combination with one or more grammars. These trained machine learning models may be trained to recognize intents, for example, by embedding recognized text from spoken utterances into a dimensionality-reduced space, and then determining which other embeddings (and therefore intents) are closest, for example, using techniques such as Euclidean distance, cosine similarity, etc. As shown in the "play <artist>" example grammar above, some grammars have slots (e.g., <artist>) that may be filled with slot values (or "parameters"). Slot values may be determined in a variety of ways. Typically, the user actively provides slot values. For example, for the grammar "Order me a <topping>pizza", the user might utter the phrase "order me a sausage pizza", in which case the slot <toppings> is automatically populated. Other slot values can be inferred based on, for example, the user location, the content currently being rendered, user preferences, and / or other cues.
[0079] A fulfillment engine (local and / or remote) may be configured to receive the predicted / estimated intent output by the NLU engine and any associated slot values and fulfill (or "parse") the intent. In various implementations, the fulfillment (or "parse") of the user intent may result in various fulfillment information (also referred to as fulfillment data), for example, generated / obtained by the fulfillment engine. This may include determining local and / or remote responses (e.g., replies) to the spoken utterance for interactions with locally installed applications performed based on the spoken utterance, commands transmitted to an Internet of Things (IoT) device (directly or via a corresponding remote system) based on the spoken utterance, and / or other parsed actions performed based on the spoken utterance. The on-device fulfillment may then initiate local and / or remote implementation / execution of the determined actions to parse the spoken utterance.
[0080] Figure 3 A flow chart illustrating an example method 300 for generating a gradient locally on a client device based on a false negative and transmitting the gradient to a remote server and / or utilizing the generated gradient to update the weights of a speech recognition model on the device is depicted. For convenience, the operations of method 300 are described with reference to a system that performs the operations. The system of method 300 includes one or more processors and / or other components of a client device. Moreover, although the operations of method 300 are shown in a particular order, this is not meant to be limiting. One or more operations may be reordered, omitted, or added.
[0081] In box 352, the system receives sensor data of one or more environmental attributes of the environment of the capture client device. In some implementations, as indicated by optional box 352A, the sensor data is non-microphone sensor data received via a non-microphone sensor. In some versions of those implementations, the sensor data includes one or more images from a camera of one or more sensor components, proximity sensor data from a proximity sensor of one or more sensor components, accelerometer data from an accelerometer of one or more sensor components, and / or magnetometer data from a magnetometer of a magnetometer of one or more sensor components. In some implementations, as indicated by optional box 352B, the sensor data is audio data that captures spoken speech and is received via one or more microphones.
[0082] At block 354, the system processes the sensor data using an on-device machine learning model to generate a prediction output. The on-device machine learning model may be, for example, a hot word detection model, a continued conversation model, a call model without hot words, and / or other machine learning models. In addition, the generated output may be, for example, a probability and / or other likelihood measure.
[0083] At box 356, the system determines whether the predicted output generated at box 354 meets the threshold. If at the iteration of box 356, the system determines that the predicted output generated at box 354 meets the threshold, the system proceeds to box 358 and initiates one or more currently dormant automated assistant functions. In some implementations, the one or more automated assistant functions include speech recognition for generating recognized text, NLU for generating natural language understanding (NLU) output, generating responses based on recognized text and / or NLU output, transmitting audio data to a remote server, and / or transmitting recognized text to a remote server. For example, assume that the predicted output generated at box 354 is a probability, and the probability must be greater than 0.85 to activate one or more currently dormant automated assistant functions, and the predicted probability is 0.88. Based on the predicted probability of 0.88 meeting the threshold of 0.85, the system proceeds to box 358 and initiates one or more currently dormant automated assistant functions as the user wishes.
[0084] If at an iteration of box 356 the system determines that the predicted output generated at box 354 fails to meet the threshold, the system proceeds to box 360 and suppresses initiation of one or more currently dormant automated assistant functions and / or shuts down one or more currently active automated assistant functions. For example, assuming that the predicted output generated at box 354 is a probability, and that the probability must be greater than 0.85 to activate one or more currently dormant automated assistant functions, the predicted probability is only 0.82. Based on the predicted probability of 0.82 failing to meet the threshold of 0.85, the system proceeds to box 360 and suppresses initiation of one or more currently dormant automated assistant functions and / or shuts down one or more currently active automated assistant functions. However, the system may perform further processing to determine whether the system should have initiated one or more currently dormant automated assistant functions in response to receiving the sensor, even though the generated predicted output failed to meet the threshold.
[0085] At frame 362, the system determines whether to receive another user interface input. In some implementations, another user interface input is to capture subsequent spoken words and another audio data received via one or more microphones. In some versions of these implementations, subsequent spoken words repeat at least a portion of the spoken words received at optional frame 352B. In other versions of those implementations, subsequent spoken words are unrelated to the spoken words received at optional frame 352B. In some implementations, another user interface input is another non-microphone sensor data received via a non-microphone sensor. In some versions of those implementations, another non-microphone sensor data includes one or more images of the camera from one or more sensor components, proximity sensor data from the proximity sensor of one or more sensor components, accelerometer data from the accelerometer of one or more sensor components and / or magnetometer data from the magnetometer of one or more sensor components. If at the iteration of frame 362, the system determines that there is no another user interface input, then the system proceeds to frame 364 and method 300 ends. If at an iteration of block 362 , the system determines that there is another user interface input received at the client device, the system proceeds to block 366 .
[0086] At box 366, the system determines whether the other user interface input received at box 362 indicates a correction of the decision made at box 356. If at an iteration of box 366, the system determines that the other user interface input received at box 362 does not indicate a correction of the decision made at box 356, the system proceeds to box 364 and method 300 ends. If at an iteration of box 366, the system determines that the other user interface input received at box 362 indicates a correction of the decision made at box 356, the system returns to box 358 and initiates one or more currently dormant automated assistant functions. Therefore, when it is determined at box 366 that there is a correction, the incorrect decision made at box 356 can be classified as an occurrence of a false negative.
[0087] Regarding hot word detection models (e.g. Figure 1B Hotword detection model 152B), assuming that the received sensor data is audio data that captures spoken utterances including the hotword "good assistant", the hotword detection model is trained to generate a prediction output indicating probability at box 354, but the probability (e.g., 0.80) fails to meet the threshold probability (e.g., 0.85) at box 356. In one instance, assuming that the other user interface input is additional audio data that captures a subsequent spoken utterance including the hotword "good assistant" that contradicts the initial decision made at box 356, the initial decision made at box 356 can be classified as incorrect (i.e., a false negative) based on, for example, determining that the duration between the spoken utterance and the additional spoken utterance satisfies a time threshold (e.g., within 3.0 seconds), a measure of similarity between the spoken utterance and the additional spoken utterance satisfies a similarity threshold (e.g., acoustic similarity of vocal characteristics, textual similarity of recognized text, and / or other similarity determinations), the magnitude of the initial probability satisfies a probability threshold about the threshold probability (e.g., within 0.20 of 0.85), a function of the duration of the initial decision at box 356 and the initial probability as described herein, and / or other determinations. In another example, assuming that the other user interface input is an alternative invocation of the assistant, such as actuation of an explicit automated assistant invocation button (e.g., a hardware button or a software button), a sensed "squeeze" of a device (e.g., invoking the assistant when the device is squeezed with at least a threshold amount of force), and / or other explicit assistant invocations that contradict the initial decision made at block 356, the initial decision made at block 356 may be classified as incorrect (i.e., a false negative) based on, for example, a determination that the duration between the spoken utterance and the alternative invocation satisfies a time threshold (e.g., within 3.0 seconds), a magnitude of the initial probability satisfies a probability threshold about a threshold probability (e.g., within 0.20 of 0.85), a function of the duration of the initial decision at block 356 and the initial probability as described herein, and / or other determinations. Thus, in these cases, the system may initiate the currently dormant automated assistant functionality based on a determination that the initial decision made at block 356 is incorrect.
[0088] Regarding the call model without hot words (for example, Figure 1C In the example of the invocation model 152C without hot words, assuming that the received sensor data is a directional gaze, mouth movement, and / or other non-speech physical movement acting as a proxy for the hot word, the invocation model without hot words is trained in box 354 to generate a prediction output indicating a probability, but the probability (e.g., 0.80) does not meet the threshold probability (e.g., 0.85) in box 356. In one example, assuming that another user interface input is audio data capturing a spoken utterance that includes the hot word "good assistant" that contradicts the initial decision made at box 356, the initial decision made at box 356 can be classified as incorrect (i.e., a false negative) based on, for example, determining that the duration between the directional gaze, mouth movement, and / or other non-speech physical movement acting as a proxy for the hot word and the spoken utterance meets a time threshold (e.g., within 3.0 seconds), the magnitude of the initial probability meets a probability threshold about the threshold probability (e.g., within 0.20 of 0.85), a function of the duration of the initial decision at box 356 and the initial probability as described herein, and / or other determinations. In another example, assuming that the other user interface input is an alternative invocation of the assistant, such as actuation of an explicit automated assistant invocation button (e.g., a hardware button or a software button), a sensed "squeeze" of a device (e.g., invoking the assistant when the device is squeezed with at least a threshold amount of force), and / or other explicit assistant invocations that contradict the initial decision made at block 356, the initial decision made at block 356 can be classified as incorrect (i.e., a false negative) based on, for example, a determination that the duration between the directed gaze, mouth movement, and / or other speechless physical movement acting as a proxy for the hotword and the alternative invocation satisfies a time threshold (e.g., within 3.0 seconds), the magnitude of the initial probability satisfies a probability threshold about the threshold probability (e.g., within 0.20 of 0.85), a function of the duration of the initial decision made at block 356 and the initial probability as described herein, and / or other determinations. Thus, in these cases, the system can initiate the currently dormant automated assistant functionality based on a determination that the initial decision made at block 356 is incorrect.
[0089] Regarding the continued conversation model (e.g. Figure 1D , assuming that the received sensor data is subsequent audio data that captures subsequent spoken utterances when the assistant has been invoked and a particular automated assistant function is active, the continued conversation model is trained at box 354 to generate a prediction output indicating a probability, but the probability (e.g., 0.80) fails to meet a threshold probability (e.g., 0.85) at box 356. In one instance, assuming that the other user interface input is additional audio data that captures an additional spoken utterance that contradicts the initial decision made at box 356 (i.e., a repetition of the subsequent spoken utterance), the initial decision made at box 356 can be classified as incorrect (i.e., a false negative) based on (for example) determining that the duration between the subsequent spoken utterance and the additional spoken utterance satisfies a time threshold (e.g., within 3.0 seconds), a measure of similarity between the subsequent spoken utterance and the additional spoken utterance satisfies a similarity threshold (e.g., acoustic similarity of vocal characteristics, textual similarity of recognized text, duration similarity between the subsequent spoken utterance and the additional spoken utterance, and / or other similarity determinations), the magnitude of the initial probability satisfies a threshold about the threshold probability (e.g., within 0.20 of 0.85), a function of the duration of the initial decision at box 356 and the initial probability as described herein, and / or other determinations. In another example, assuming that the other user interface input is an alternative invocation of the assistant, such as actuation of an explicit automated assistant invocation button (e.g., a hardware button or a software button), a sensed "squeeze" of a device (e.g., invoking the assistant when the device is squeezed with at least a threshold amount of force), and / or other explicit assistant invocations that contradict the initial decision made at block 356, the initial decision made at block 356 can be classified as incorrect (i.e., a false negative) based on, for example, a determination that the duration between the subsequent spoken utterance and the alternative invocation satisfies a time threshold (e.g., within 3.0 seconds), a magnitude of the initial probability satisfies a probability threshold about a threshold probability (e.g., within 0.20 of 0.85), a function of the duration of the initial decision at block 356 and the initial probability as described herein, and / or other determinations. Thus, in these cases, the system can refrain from closing the currently active automated assistant function (i.e., because the assistant has already been invoked) and / or initiating another currently dormant automated assistant function based on a determination that the initial decision made at block 356 is incorrect.
[0090] Additionally, if at an iteration of box 366 , the system determines that another user interface input received at box 362 indicates a correction to the decision made at box 356 , the system provides the prediction output generated at box 354 to box 368 .
[0091] At block 368, the system generates a gradient based on comparing the predicted output to a ground truth output that satisfies a threshold. In some implementations, the ground truth output corresponds to the output that satisfies the threshold at block 356 and is generated based on determining at block 366 another user interface input received at block 362 indicating a correction of the decision made at block 356. For example, for a false negative, if the generated predicted output is 0.82 and the threshold is 0.85, the system can generate a ground truth output of 1.0. In such an example, generating a gradient is based on comparing a predicted output of 0.82 to a ground truth output of 0.1.
[0092] At box 370, the system updates one or more weights of the machine learning model on the device based on the generated gradient, and / or the system transmits the generated gradient to a remote system (e.g., via the Internet or other wide area network) (without transmitting any of the audio data, sensor data, and / or another user interface input). When the gradient is transmitted to the remote system, the remote system uses the generated gradient and the additional gradient from the additional client device to update the global weights of the global speech recognition model. After box 370, the system then proceeds back to box 352.
[0093] Figure 4 A flow chart illustrating an example method 400 for generating a gradient locally on a client device based on a false positive and transmitting the gradient and / or utilizing the generated gradient to update the weights of a speech recognition model on the device is depicted. For convenience, the operation of method 400 is described with reference to a system that performs the operation. The system of method 400 includes one or more processors and / or other components of a client device. In addition, although the operation of method 400 is shown in a particular order, this is not meant to be limiting. One or more operations may be reordered, omitted, or added.
[0094] At box 452, the system receives sensor data of one or more environmental attributes of the environment of the capture client device. In some implementations, as indicated by optional box 452A, the sensor data is non-microphone sensor data received via a non-microphone sensor. In some versions of those implementations, the sensor data includes one or more images of a camera from one or more sensor components, proximity sensor data from a proximity sensor of one or more sensor components, accelerometer data from an accelerometer of one or more sensor components, and / or magnetometer data from a magnetometer of a magnetometer of one or more sensor components. In some implementations, as indicated by optional box 452B, the sensor data is audio data that captures spoken speech and is received via one or more microphones.
[0095] At block 454, the system processes the sensor data using an on-device machine learning model to generate a prediction output. The on-device machine learning model may be, for example, a hot word detection model, a continued conversation model, a call model without hot words, and / or other machine learning models. In addition, the generated output may be, for example, a probability and / or other likelihood measure.
[0096] At box 456, the system determines whether the predicted output generated at box 454 meets the threshold. If at the iteration of box 356, the system determines that the predicted output generated at box 354 fails to meet the threshold, the system proceeds to box 458 and suppresses the initiation of one or more currently dormant automated assistant functions and / or closes one or more currently active automated assistant functions. In some implementations, the one or more automated assistant functions include speech recognition for generating recognized text, NLU for generating natural language understanding (NLU) output, generating responses based on recognized text and / or NLU output, transmitting audio data to a remote server, and / or transmitting recognized text to a remote server. For example, assume that the predicted output generated at box 454 is a probability, and the probability must be greater than 0.85 to activate one or more currently dormant automated assistant functions, but the predicted probability is only 0.82. Based on the predicted probability of 0.82 failing to meet the threshold of 0.85, the system proceeds to box 458 and suppresses the initiation of one or more currently dormant automated assistant functions and / or closes one or more currently active automated assistant functions.
[0097] If, at an iteration of box 456, the system determines that the predicted output generated at box 454 satisfies the threshold, the system proceeds to box 460 and initiates one or more currently dormant automated assistant functions. For example, assume that the predicted output generated at box 454 is a probability, and the probability must be greater than 0.85 to activate one or more currently dormant automated assistant functions, and the predicted probability is 0.88. Based on the predicted probability of 0.88 satisfying the threshold of 0.85, the system proceeds to box 460 and initiates one or more currently dormant automated assistant functions as the user wishes. However, although the generated predicted output satisfies the threshold, the system may perform another process in response to receiving the sensor to determine whether the system should have suppressed initiation of one or more currently dormant automated assistant functions and / or should have shut down one or more currently active automated assistant functions.
[0098] In box 462, the system determines whether another user interface input is received. In some implementations, another user interface input is another non-microphone sensor data received via a non-microphone sensor. In some versions of those implementations, another non-microphone sensor data includes one or more images of a camera from one or more sensor components, proximity sensor data from a proximity sensor of one or more sensor components, accelerometer data from an accelerometer of an accelerometer of one or more sensor components, and / or magnetometer data from a magnetometer of a magnetometer of one or more sensor components. In some implementations, another user interface input is another audio data that captures subsequent spoken words and is received via one or more microphones. In some versions of those implementations, subsequent spoken words repeat at least a portion of the spoken words received at optional box 452B. In other versions of those implementations, subsequent spoken words are not related to the spoken words received at optional box 452B. If at the iteration of box 462, the system determines that there is no other user interface input, the system proceeds to box 464 and method 400 ends. If at the iteration of box 462, the system determines that there is another user interface input received at the client device, the system proceeds to box 466.
[0099] At box 466, the system determines whether another user interface input received at box 462 indicates a correction of the decision made at box 456. If at an iteration of box 466, the system determines that another user interface input received at box 462 does not indicate a correction of the decision made at box 456, the system proceeds to box 464 and method 400 ends. If at an iteration of box 466, the system determines that another user interface input received at box 462 indicates a correction of the decision made at box 456, the system returns to box 458 and suppresses initiation of one or more currently dormant automated assistant functions and / or closes one or more currently active automated assistant functions. Therefore, when it is determined at box 466 that a correction exists, the incorrect decision made at box 456 can be classified as a false positive occurrence.
[0100] Regarding hot word detection models (e.g. Figure 1B ), assuming that the received sensor data is audio data that captures a spoken utterance including "show consistent" as opposed to the hot word "good assistant", the hot word detection model is trained to generate a prediction output indicating a probability at box 454, and the probability (e.g., 0.90) meets a threshold probability (e.g., 0.85) at box 456. In one instance, assuming that the other user interface input is additional audio data that captures subsequent spoken utterances including "no", "stop", "cancel", and / or another spoken utterance that contradicts the initial decision made at box 456, the initial decision made at box 456 can be classified as incorrect (i.e., a false positive) based on, for example, determining that the duration between the spoken utterance and the additional spoken utterance meets a time threshold (e.g., within 3.0 seconds), the magnitude of the initial probability meets a probability threshold about the threshold probability (e.g., within 0.20 of 0.85), a function of the duration of the initial decision at box 456 and the initial probability as described herein, and / or other determinations. In another example, assuming that the other user interface input is an alternative input cancel invocation of the assistant, such as actuation of an explicit automated assistant invocation button (e.g., a hardware button or a software button), a sensed "squeeze" of the device (e.g., invoking the assistant when the device is squeezed with at least a threshold amount of force), and / or other explicit input cancel invocation of the assistant, the initial decision made at box 456 can be classified as incorrect (i.e., a false positive) based on, for example, a determination that the duration between the spoken utterance and the alternative input cancel invocation satisfies a time threshold (e.g., within 3.0 seconds), the magnitude of the initial probability satisfies a probability threshold about the threshold probability (e.g., within 0.20 of 0.85), a function of the duration and initial probability of the initial decision at box 456 as described herein, and / or other determinations. Thus, in these cases, the system can refrain from initiating a currently dormant automated assistant function and / or shut down a currently active automated assistant function based on a determination that the initial decision made at box 456 is incorrect.
[0101] Regarding the call model without hot words (for example, Figure 1C In the example of the invocation model 152C of the non-hotword, assuming that the received sensor data is a directional gaze, mouth movement, and / or other non-speech physical movement acting as a proxy for the hotword, the non-hotword invocation model is trained to generate a prediction output indicating a probability at box 454, and the probability (e.g., 0.90) meets a threshold probability (e.g., 0.85) at box 456. In one example, assuming that another user interface input is audio data capturing a spoken utterance including "no", "stop", "cancel", and / or another spoken utterance that contradicts the initial decision made at box 456, the initial decision made at box 456 can be classified as incorrect (i.e., a false positive) based on, for example, determining that the duration between the directional gaze, mouth movement, and / or other non-speech physical movement acting as a proxy for the hotword and the spoken utterance meets a time threshold (e.g., within 3.0 seconds), the magnitude of the initial probability meets a probability threshold about the threshold probability (e.g., within 0.20 of 0.85), and / or other determinations. In another example, assuming that another user interface input is another sensor data that negates the assistant's directional gaze and / or alternative input cancel invocation, such as actuation of an explicit automated assistant invocation button (e.g., a hardware button or a software button), a sensed "squeeze" of the device (e.g., invoking the assistant when the device is squeezed with at least a threshold amount of force), and / or other explicit input cancel invocation of the assistant, the initial decision made at box 456 can be classified as incorrect (i.e., a false positive) based on, for example, a determination that the duration between the directional gaze, mouth movement, and / or other speechless physical movement acting as a proxy for the hotword and the assistant's alternative input cancel invocation satisfies a time threshold (e.g., within 3.0 seconds), the magnitude of the initial probability satisfies a probability threshold about the threshold probability (e.g., within 0.20 of 0.85), a function of the duration of the initial decision at box 456 and the initial probability as described herein, and / or other determinations. Thus, in these cases, the system can refrain from initiating a currently dormant automated assistant function and / or shutting down a currently active automated assistant function based on a determination that the initial decision made at box 456 is incorrect.
[0102] Regarding the continued conversation model (e.g. Figure 1D 454 ), assuming that the received sensor data is subsequent audio data that captures a subsequent spoken utterance while the assistant has been invoked and a particular automated assistant function is active, the continued conversation model is trained to generate a predicted output at box 454 indicating a probability, and that the probability (e.g., 0.90) satisfies a threshold probability (e.g., 0.85) at box 456. In one example, assuming that another user interface input is additional audio data that captures an additional spoken utterance including "no," "stop," "cancel," and / or another spoken utterance that contradicts the initial decision made at box 456, the initial decision made at box 456 can be classified as incorrect (i.e., a false positive) based on, for example, a determination that a duration between the subsequent spoken utterance and the additional spoken utterance satisfies a time threshold (e.g., within 3.0 seconds), a magnitude of the initial probability satisfies a probability threshold about the threshold probability (e.g., within 0.20 of 0.85), a function of the duration of the initial decision at box 456 and the initial probability as described herein, and / or other determinations. In another example, assuming that the other user interface input is an alternative input cancel invocation of the assistant, such as actuation of an explicit automated assistant invocation button (e.g., a hardware button or a software button), a sensed device "squeeze" (e.g., invoking the assistant when the device is squeezed with at least a threshold amount of force), and / or other explicit input cancel invocation of the assistant, the initial decision made at box 456 can be classified as incorrect (i.e., a false positive) based on, for example, a determination that the duration between the subsequent spoken utterance and the alternative invocation satisfies a time threshold (e.g., within 3.0 seconds), the magnitude of the initial probability satisfies a probability threshold about the threshold probability (e.g., within 0.20 of 0.85), a function of the duration and initial probability of the initial decision at box 456 as described herein, and / or other determinations. Thus, in these examples, the system can shut down the currently active automated assistant function (i.e., because the assistant has already been invoked) and / or suppress initiation of the currently dormant automated assistant function based on a determination that the initial decision made at box 456 is incorrect.
[0103] Additionally, if at an iteration of box 466 , the system determines that another user interface input received at box 462 indicates a correction to the decision made at box 456 , the system provides the prediction output generated at box 454 to box 468 .
[0104] At block 468, the system generates a gradient based on comparing the predicted output to the ground truth output that meets the threshold. In some implementations, the ground truth output corresponds to the output that failed to meet the threshold at block 456, and is generated based on determining at block 466 that another user interface input received at block 462 indicates a correction of the decision made at block 456. For example, for a false positive, if the generated predicted output is 0.88 and the threshold is 0.85, the system can generate a ground truth output of 0.0. In this example, generating the gradient is based on comparing the predicted output of 0.88 to the ground truth output of 0.0.
[0105] At box 470, the system updates one or more weights of the machine learning model on the device based on the generated gradient, and / or the system transmits the generated gradient to a remote system (e.g., via the Internet or other wide area network) (without transmitting any of the audio data, sensor data, and / or another user interface input). When the gradient is transmitted to the remote system, the remote system uses the generated gradient and the additional gradient from the additional client device to update the global weights of the global speech recognition model. After block 470, the system then proceeds back to box 452.
[0106] Note that in various implementations of methods 300 and 400, the audio data, the predicted output, another user interface input, and / or the ground truth output can be stored locally on the client device. In addition, in some versions of those implementations of methods 300 and 400, in response to determining that the current state of the client device satisfies one or more conditions, generating gradients, updating one or more weights of the machine learning model on the device, and / or transmitting the gradients to the remote system are performed. For example, one or more conditions include that the client device is charging, the client device has at least a threshold charging state, and / or the client device is not carried by the user. Moreover, in some additional or alternative versions of those implementations of methods 300 and 400, generating gradients, updating one or more weights of the machine learning model on the device, and / or transmitting the gradients to the remote system are performed in real time. In these and other ways, the machine learning model on the device can be quickly adapted to mitigate the occurrence of false negatives and / or false positives. In addition, this enables the performance of the machine learning model on the device to be improved for the attributes of the user of the client device (such as the pitch, intonation, accent, and / or other voice characteristics in the case of a machine learning model on a device that processes audio data that captures spoken utterances).
[0107] Figure 5 A flow chart illustrating an example method 500 for updating the weights of a global speech recognition model based on gradients received from a remote client device and transmitting the updated weights or updated global speech recognition model to the remote client device is depicted. For convenience, the operations of method 500 are described with reference to a system that performs the operations. The system may include various components of various computer systems, such as one or more server devices. Moreover, although the operations of method 500 are shown in a particular order, this is not meant to be limiting. One or more operations may be reordered, omitted, or added.
[0108] At block 552, the system receives a gradient from a remote client device. For example, the system may receive a gradient from a remote client device that is executing Figure 3 The corresponding example of method 300 and / or Figure 4 In an example of method 400 , a plurality of client devices receive a gradient.
[0109] At block 554, the system updates the weights of the global speech recognition model based on the gradients received at block 552. Iterations of blocks 552 and 554 may continue to be performed as new gradients are received and / or queued after being received.
[0110] In box 556, the system at least periodically determines whether one or more conditions, such as one or more conditions described herein, are met. Typically, the condition acts as an agent for determining whether the global model has been updated to the extent that the utilization of network resources is justified when the updated weights of the transmission model and / or the updated model itself are transmitted. In other words, the condition is used as an agent for determining whether the performance gain of the model proves the reasonable use of network resources. If so, the system proceeds to box 558, and transmits the currently updated weights and / or the currently updated global speech recognition model to multiple client devices. The updated weights and / or global speech recognition model can be optionally transmitted to a given client device in response to a request from a given client device (such as a request during the update process and / or a request sent due to the client device being idle and / or charging).
[0111] Figure 6 610 is a block diagram of an example computing device 610 that can optionally be used to perform one or more aspects of the techniques described herein. In some implementations, one or more of the client device, cloud-based automated assistant components, and / or other components can include one or more components of the example computing device 610.
[0112] The computing device 610 typically includes at least one processor 614 that communicates with a number of peripheral devices via a bus subsystem 612. These peripheral devices may include: a storage subsystem 624, including, for example, a memory subsystem 625 and a file storage subsystem 626; a user interface output device 620; a user interface input device 622; and a network interface subsystem 616. The input and output devices allow a user to interact with the computing device 610. The network interface subsystem 616 provides an interface to an external network and is coupled to corresponding interface devices in other computing devices.
[0113] User interface input devices 622 may include: a keyboard; a pointing device such as a mouse, trackball, touchpad, or graphics tablet; a scanner; a touch screen incorporated into a display; an audio input device such as a voice recognition system, a microphone; and / or other types of input devices. In general, the use of the term "input device" is intended to include all possible types of devices and ways of inputting information into computing device 610 or a communication network.
[0114] User interface output device 620 can include a display subsystem, a printer, a fax machine, or a non-visual display such as an audio output device. The display subsystem can include a cathode ray tube (CRT), a flat panel device such as a liquid crystal display (LCD), a projection device, or some other mechanism for creating a visible image. The display subsystem can also provide a non-visual display such as via an audio output device. In general, the use of the term "output device" is intended to include all possible types of devices and the manner in which information is output from computing device 610 to a user or another machine or computing device.
[0115] The storage subsystem 624 stores programming and data structures that provide the functionality of some or all of the modules described herein. For example, the storage subsystem 624 may include a computer program for performing selected aspects of the methods disclosed herein and implementing the Figure 1A and 1B The logic of the various components is depicted in .
[0116] These software modules are typically executed by the processor 614 alone or in combination with other processors. The memory 625 used in the storage subsystem 624 may include multiple memories, including a main random access memory (RAM) 630 for storing instructions and data during program execution and a read-only memory (ROM) 632 in which fixed instructions are stored. The file storage subsystem 626 may provide persistent storage for program and data files, and may include a hard drive, a floppy disk drive and associated removable media, a CD-ROM drive, an optical drive, or a removable media cartridge. Modules that implement the functionality of certain implementations may be stored in the storage subsystem 624 by the file storage subsystem 626, or in other machines accessible to the processor 614.
[0117] The bus subsystem 612 provides a mechanism for the various components and subsystems of the computing device 610 to communicate with each other as intended. Although the bus subsystem 612 is shown schematically as a single bus, alternative implementations of the bus subsystem may use multiple busses.
[0118] The computing device 610 can be of various types, including a workstation, a server, a computing cluster, a blade server, a server farm, or any other data processing system or computing device. Due to the ever-changing nature of computers and networks, Figure 6 The description of computing device 610 depicted in FIG. 6 is intended only as a specific example for purposes of illustrating some implementations. Many other configurations of computing device 610 may have more Figure 6 The computing device depicted may have more or fewer components.
[0119] In situations where the systems described herein collect or otherwise monitor personal information about a user, or where personal and / or monitored information may be exploited, the user may be provided with the opportunity to control whether a program or feature collects user information (e.g., information about the user's social network, social actions or activities, occupation, user preferences, or the user's current geographic location), or to control whether and / or how content that may be more relevant to the user is received from a content server. In addition, certain data may be processed in one or more ways before being stored or used to remove personally identifiable information. For example, the user's identity may be processed so that the user's personally identifiable information cannot be determined, or the user's geographic location may be generalized (such as to a city, zip code, or state level) when obtaining the location of geographic location information so that the user's specific geographic location cannot be determined. Thus, the user may have control over how information about the user is collected and / or used.
[0120] In some implementations, a method performed by one or more processors of a client device is provided and includes receiving audio data that captures a user's spoken utterances via one or more microphones of the client device. The method also includes processing the audio data using a machine learning model stored locally at the client device to generate a predicted output. The method also includes making a decision to suppress the initiation of one or more currently dormant automated assistant functions based on the predicted output failing to meet a threshold. The method also includes, after making the decision to suppress the initiation of the one or more currently dormant automated assistant functions, determining that the decision is incorrect based on another user interface input received at the client device after receiving the audio data. The method also includes, in response to determining that the decision is incorrect, generating a gradient based on comparing the predicted output to a ground truth output that meets the threshold, and updating one or more weights of the machine learning model based on the generated gradient.
[0121] These and other implementations of the technology may include one or more of the following features.
[0122] In some implementations, determining that the decision is incorrect is further based on a size of the predicted output. In some versions of those implementations, determining that the decision is incorrect further based on a size of the predicted output includes determining that the predicted output is within a threshold range of a threshold for initiating the one or more currently dormant automated assistant functions while failing to meet the threshold.
[0123] In some implementations, determining that the decision is incorrect based on another user interface input received at the client device after receiving the audio data is based on a duration between receiving the audio data and receiving the another user interface input.
[0124] In some implementations, the other user interface input is an additional spoken utterance captured in additional audio data. In some versions of those implementations, the method further includes processing the additional audio data using the machine learning model to generate an additional prediction output, and making an additional decision to initiate the one or more currently dormant automated assistant functions based on the additional prediction output satisfying the threshold. Determining that the decision is incorrect based on another user interface input received at the client device after receiving the audio data includes determining that the decision is incorrect based on the additional decision to initiate the one or more currently dormant automated assistant functions.
[0125] In some implementations, the other user interface input is an additional spoken utterance captured in the additional audio data. In some versions of those implementations, the method further includes determining one or more metrics of similarity between the spoken utterance and the additional spoken utterance. Determining that the decision is incorrect based on another user interface input received at the client device after receiving the audio data is based on one or more metrics of similarity between the spoken utterance and the additional spoken utterance. In some versions of those implementations, the one or more metrics of similarity include duration similarity based on a comparison of the duration of the spoken utterance with the duration of the additional spoken utterance, voice similarity based on a comparison of the voice characteristics of the spoken utterance with the voice characteristics of the additional spoken utterance, and / or text similarity based on a comparison of the recognized text of the spoken utterance and the recognized text of the additional spoken utterance.
[0126] In some implementations, the other user interface input is an additional spoken utterance captured in additional audio data. In some versions of those implementations, determining that the decision is incorrect based on another user interface input received at the client device after receiving the audio data includes determining that the decision is incorrect based on one or more acoustic features of the additional spoken utterance, or text recognized from the additional spoken utterance using a speech recognition model stored locally at the client device.
[0127] In some implementations, determining that the decision is incorrect includes determining a confidence metric indicating a confidence that the decision is incorrect. In some versions of those implementations, the method also includes determining a magnitude of the ground truth output that satisfies the threshold based on the confidence metric.
[0128] In some implementations, the one or more currently dormant automated assistant functions include speech recognition, natural language understanding (NLU), transmitting the audio data or subsequent audio data to a remote server, transmitting recognized text from the speech recognition to a remote server, and / or generating a response based on the recognized text and / or NLU output from the NLU.
[0129] In some implementations, the machine learning model is a hot word detection model. In some versions of those implementations, the one or more currently dormant automated assistant functions include speech recognition using a speech recognition model stored locally at the client device, transmitting the audio data to a remote server, transmitting recognized text from the speech recognition to the remote server, and / or performing natural language understanding on the recognized text using a natural language understanding model stored locally at the client device.
[0130] In some implementations, the machine learning model is a continued conversation model. In some versions of those implementations, the one or more currently dormant automated assistant functions include transmitting the audio data to a remote server, transmitting recognized text from a local speech recognition of the audio data to the remote server, and / or generating a response based on the audio data or the recognized text. In some other versions of those implementations, the predicted output is also based on natural language understanding data generated using the machine learning model to process the recognized text and / or based on the recognized text.
[0131] In some implementations, the method further includes transmitting the generated gradient to a remote system over a network without transmitting any of the following: the audio data and the other user interface input. The remote system updates the global weights of a global machine learning model corresponding to the machine learning model using the generated gradient and the additional gradient from the additional client device. In some versions of those implementations, the updated global weights of the global speech recognition model are stored in a memory of the remote system. In some versions of those implementations, the method further includes receiving the global machine learning model from the remote system at the client device. Receiving the global machine learning model is after the remote system updates the global weights of the global machine learning model based on the generated gradient and the additional gradient. In some versions of those implementations, the method further includes replacing the machine learning model with the global machine learning model in a local storage of the client device in response to receiving the global machine learning model. In some versions of those implementations, the method further includes receiving the updated global weights from the remote system at the client device. Receiving the updated global weights is after the remote system updates the global weights of the global machine learning model based on the generated gradient and the additional gradient. In some versions of those implementations, the method further includes, in response to receiving the updated global weights, replacing weights of the machine learning model with the updated global weights in local storage of the client device.
[0132] In some implementations, the method further includes determining that a current state of the client device satisfies one or more conditions based on sensor data from one or more sensors of the client device. In response to determining that the current state of the client device satisfies the one or more conditions, generating the gradient and / or updating the one or more weights is performed.
[0133] In some implementations, a method performed by one or more processors of a client device is provided and includes receiving sensor data that captures one or more environmental attributes of an environment of the client device via one or more sensor components of the client device. The method also includes processing the sensor data using a machine learning model stored locally at the client device to generate a prediction output that specifies whether one or more currently dormant automated assistant functions are activated. The method also includes making a decision about whether to trigger the one or more currently dormant automated assistant functions based on the predicted output failing to meet a threshold. The method also includes, after making the decision, determining that the decision is incorrect. The method also includes, in response to determining that the determination is incorrect, generating a gradient based on comparing the predicted output to a ground truth output that meets the threshold, and updating one or more weights of the machine learning model based on the generated gradient.
[0134] These and other implementations of the technology may include one or more of the following features.
[0135] In some implementations, the machine learning model is a hot word-free call model. In some versions of those implementations, the sensor data includes one or more images from a camera of the one or more sensor components, proximity sensor data from a proximity sensor of the one or more sensor components, accelerometer data from an accelerometer of the one or more sensor components, and / or magnetometer data from a magnetometer of the one or more sensor components.
[0136] In some implementations, the one or more currently dormant automated assistant functions include speech recognition using a speech recognition model stored locally at the client device, transmitting the audio data to a remote server, transmitting recognized text from the speech recognition to the remote server, and / or performing natural language understanding on the recognized text using a natural language understanding model stored locally at the client device.
[0137] In some implementations, determining that the decision is incorrect includes receiving additional user interface input at the client device after receiving the sensor data, and determining that the additional user interface input indicates a correction to the decision. Determining that the decision is incorrect is based on determining that the additional user interface input indicates a correction to the decision.
[0138] In some implementations, determining that the additional user interface input indicates that the correction to the determination is based on the duration between receiving the sensor data and receiving the additional user interface input. In some versions of those implementations, the sensor data includes audio data that captures the spoken utterance, and the additional user interface input is an additional spoken utterance captured in the additional audio data. In some other versions of those implementations, the method also includes determining one or more measures of similarity between the spoken utterance and the additional spoken utterance based on the audio data and the additional audio data. Determining that the additional user interface input indicates that the correction to the determination is based on one or more measures of similarity. In some versions of those implementations, the additional user interface input is additional audio data, and determining that the additional user interface input indicates that the correction to the decision is based on one or more acoustic features of the additional audio data and / or text recognized from the additional audio data using a speech recognition model stored locally at the client device.
[0139] In some implementations, determining that the decision is incorrect is also based on a magnitude of the predicted output.
[0140] In some implementations, the decision does not trigger the one or more currently dormant automated assistant functions. In some versions of those implementations, determining that the decision is incorrect is based on processing the additional user interface input using the machine learning model to generate additional predicted outputs, and determining to trigger the one or more currently dormant automated assistant functions based on the additional predicted outputs.
[0141] In some implementations, the decision is to trigger the one or more currently dormant automated assistant functions. In some versions of those implementations, the one or more currently dormant automated assistant functions that are triggered include transmitting the audio data to a remote server. In some other versions of those implementations, determining that the decision is incorrect includes receiving, in response to the transmitting, an indication from the remote server that the determination is incorrect.
[0142] In some implementations, a method performed by one or more processors of a client device is provided and includes receiving audio data capturing a user's spoken utterance via one or more microphones of the client device. The method also includes processing the audio data using a machine learning model stored locally at the client device to generate a predicted output. The method also includes making a decision to suppress initiation of one or more currently dormant automated assistant functions based on the predicted output failing to meet a threshold. The method also includes, after making a decision to suppress initiation of the one or more currently dormant automated assistant functions, determining that the decision is incorrect based on another user interface input received at the client device after receiving the audio data. The method also includes, in response to determining that the decision is incorrect, generating a gradient based on comparing the predicted output with a ground truth output that meets the threshold, and transmitting the generated gradient to a remote system over a network without transmitting the audio data and / or the other user interface input. The remote system updates the global weights of a global speech recognition model using the generated gradient and additional gradients from additional client devices.
[0143] These and other implementations of the technology may include one or more of the following features.
[0144] In some implementations, the updated global weights of the global speech recognition model are stored in a memory of the remote system.
[0145] In some implementations, the method further includes receiving the global speech recognition model from the remote system at the client device. Receiving the global speech recognition model is after the remote system updates the global weights of the global speech recognition model based on the generated gradients and the additional gradients. In some implementations, the method further includes replacing the speech recognition model with the global speech recognition model in the local storage of the client device in response to receiving the global speech recognition model.
[0146] In some implementations, the method further includes receiving the updated global weights from the remote system at the client device. Receiving the updated global weights is after the remote system updates the global weights of the global end-to-end speech recognition model based on the gradients and the additional gradients. In some implementations, the method further includes replacing the weights of the speech recognition model with the updated global weights in the local storage of the client device in response to receiving the updated global weights.
[0147] In some implementations, a method performed by one or more processors of a client device is provided and includes receiving sensor data capturing one or more environmental attributes of an environment of the client device via one or more sensor components of the client device. The method also includes processing the sensor data using a machine learning model stored locally at the client device to generate a prediction output specifying whether one or more currently dormant automated assistant functions are activated. The method also includes making a decision about whether to trigger the one or more currently dormant automated assistant functions based on the predicted output failing to meet a threshold. The method also includes, after making the decision, determining that the decision is incorrect. The method also includes, in response to determining that the determination is incorrect, generating a gradient based on comparing the predicted output with a ground truth output that meets the threshold, and transmitting the generated gradient to a remote system over a network without transmitting the audio data and / or the other user interface input. The remote system updates the global weights of a global speech recognition model using the generated gradient and additional gradients from additional client devices.
[0148] These and other implementations of the technology may include one or more of the following features.
[0149] In some implementations, the updated global weights of the global speech recognition model are stored in a memory of the remote system.
[0150] In some implementations, the method further includes receiving the global speech recognition model from the remote system at the client device. Receiving the global speech recognition model is after the remote system updates the global weights of the global speech recognition model based on the generated gradients and the additional gradients. In some implementations, the method further includes replacing the speech recognition model with the global speech recognition model in the local storage of the client device in response to receiving the global speech recognition model.
[0151] In some implementations, the method further includes receiving the updated global weights from the remote system at the client device. Receiving the updated global weights is after the remote system updates the global weights of the global end-to-end speech recognition model based on the gradients and the additional gradients. In some implementations, the method further includes replacing the weights of the speech recognition model with the updated global weights in the local storage of the client device in response to receiving the updated global weights.< / topping> < / artist>
Claims
1. A method performed by one or more processors of a client device, the method comprising: receiving, via one or more microphones of the client device, audio data capturing spoken utterances of a user; processing the audio data using a machine learning model stored locally at the client device to generate a predicted output; making a determination to suppress initiation of one or more currently dormant automated assistant functions based on the predicted output failing to satisfy a threshold; Upon making a decision to suppress initiation of the one or more currently dormant automated assistant functions: receiving, via the client device, another user interface input to initiate one or more of the one or more currently dormant automated assistant functions and contradicting a decision to refrain from initiating one or more of the one or more currently dormant automated assistant functions; and determining that the decision to suppress initiation of one or more of the one or more currently dormant automated assistant functions is incorrect based on a duration between receiving the audio data processed in making the decision and receiving the other user interface input that contradicts the decision; as well as In response to determining that the decision is incorrect: generating a gradient based on comparing the predicted output to a ground truth output that satisfies the threshold, and One or more of the following: updating one or more weights of the machine learning model based on the generated gradients; or transmitting the generated gradient to a remote system over a network without transmitting any of: the audio data and the another user interface input; The remote system uses the generated gradients and additional gradients from additional client devices to update the global weights of a global machine learning model corresponding to the machine learning model.
2. The method according to claim 1, wherein: Determining that the decision is incorrect is further based on a magnitude of the predicted output.
3. The method according to claim 2, wherein: Determining that the determination is incorrect further based on the magnitude of the predicted output includes determining that the predicted output is within a threshold range of a threshold for initiating the one or more currently dormant automated assistant functions while failing to satisfy the threshold.
4. The method according to claim 1, wherein: The further user interface input is an additional spoken utterance captured in additional audio data, and further comprising: processing the additional audio data using the machine learning model to generate additional predicted outputs; and making an additional decision to initiate the one or more currently dormant automated assistant functions based on the additional predicted output satisfying the threshold; Wherein, based on another user interface input received at the client device after receiving the audio data, determining that the decision is incorrect comprises: The decision is determined to be incorrect based on an additional decision to initiate the one or more currently dormant automated assistant functions.
5. The method according to claim 1, wherein: The further user interface input is an additional spoken utterance captured in additional audio data, and further comprising: determining one or more measures of similarity between the spoken utterance and the additional spoken utterance; Wherein determining that the determination is incorrect based on another user interface input received at the client device after receiving the audio data is based on one or more measures of similarity between the spoken utterance and the additional spoken utterance.
6. The method according to claim 5, wherein: The one or more measures of similarity include one or more of the following: a duration similarity based on a comparison of the duration of the spoken utterance to the duration of the additional spoken utterance, vocal similarity based on a comparison of vocal characteristics of the spoken utterance with vocal characteristics of the additional spoken utterance, and A text similarity based on a comparison of the recognized text of the spoken utterance and the recognized text of the additional spoken utterance.
7. The method according to claim 1, wherein: The other user interface input is an additional spoken utterance captured in additional audio data, and wherein determining that the decision is incorrect based on another user interface input received at the client device after receiving the audio data comprises: The determination is incorrect based on: one or more acoustic features of the additional spoken utterance, or Text is recognized from the additional spoken utterance using a speech recognition model stored locally at the client device.
8. The method according to claim 1, wherein: The one or more currently dormant automated assistant functions include one or more of the following: Speech recognition, Natural Language Understanding (NLU), transmitting the audio data or subsequent audio data to a remote server, transmitting the recognized text from the speech recognition to a remote server, and A response is generated based on the recognized text and / or an NLU output from the NLU.
9. The method according to claim 1, wherein: The machine learning model is a hotword detection model, and wherein the one or more currently dormant automated assistant functions include one or more of the following: speech recognition using a speech recognition model stored locally at the client device, transmitting the audio data to a remote server, transmitting the recognized text from the speech recognition to the remote server, and The recognized text is subjected to natural language understanding using a natural language understanding model stored locally at the client device.
10. The method according to claim 1, wherein: The machine learning model is a continued conversation model, and wherein the one or more currently dormant automated assistant functions include one or more of the following: transmitting the audio data to a remote server, transmitting recognized text from the local speech recognition of the audio data to the remote server, and A response is generated based on the audio data or the recognized text.
11. The method according to claim 1, further comprising: receiving, at the client device, updated global weights from the remote system, wherein receiving the updated global weights is after the remote system updates the global weights of the global machine learning model based on the generated gradients and the additional gradients; and In response to receiving the updated global weights, replacing the weights of the machine learning model with the updated global weights in local storage of the client device.
12. The method according to any one of claims 1 to 11, further comprising: determining, based on sensor data from one or more sensors of the client device, that a current state of the client device satisfies one or more conditions, Wherein, generating the gradient and / or updating the one or more weights is performed in response to determining that the current state of the client device satisfies the one or more conditions.
13. A method performed by one or more processors of a client device, the method comprising: receiving, via one or more sensor components of the client device, sensor data capturing one or more environmental attributes of an environment of the client device; processing the sensor data using a machine learning model stored locally at the client device to generate a prediction output specifying whether one or more currently dormant automated assistant functions are to be activated; making a determination as to whether to trigger the one or more currently dormant automated assistant functions based on the predicted output failing to satisfy a threshold; After said decision has been made: receiving, via the client device, another user interface input to initiate one or more of the one or more currently dormant automated assistant functions and contradicting a decision to refrain from initiating one or more of the one or more currently dormant automated assistant functions; and determining that the decision to suppress initiation of one or more of the one or more currently dormant automated assistant functions is incorrect based on a duration between receiving the sensor data processed in making the decision and receiving the other user interface input that contradicts the decision; as well as In response to determining that the determination is incorrect: generating a gradient based on comparing the predicted output to a ground truth output that satisfies the threshold, and One or more of the following: updating one or more weights of the machine learning model based on the generated gradients; or transmitting the generated gradient to a remote system over a network without transmitting any of: the sensor data and the other user interface input; The remote system uses the generated gradients and additional gradients from additional client devices to update the global weights of a global machine learning model corresponding to the machine learning model.
14. The method according to claim 13, wherein: The machine learning model is a hot word-free call model, and wherein the sensor data includes: one or more images from a camera in the one or more sensor components, proximity sensor data from a proximity sensor in the one or more sensor components, accelerometer data from an accelerometer in the one or more sensor components, and / or magnetometer data from a magnetometer in the one or more sensor components.
15. The method according to claim 13, wherein: The one or more currently dormant automated assistant functions include one or more of the following: speech recognition using a speech recognition model stored locally at the client device, transmitting the sensor data to a remote server, transmitting the recognized text from the speech recognition to the remote server, and The recognized text is subjected to natural language understanding using a natural language understanding model stored locally at the client device.
16. The method according to any one of claims 13 to 15, wherein: The decision does not trigger the one or more currently dormant automated assistant functions, and wherein determining that the decision is incorrect is based on: processing the another user interface input using the machine learning model to generate an additional predicted output; as well as Determining triggering the one or more currently dormant automated assistant functions is performed based on the additional predicted output.
17. A client device comprising: At least one microphone; at least one display; as well as One or more processors, the one or more processors executing locally stored instructions to cause the processors to perform the method according to any one of claims 1 to 16.
18. A computer program product comprising instructions which, when executed by one or more processors, cause the one or more processors to perform the method according to any one of claims 1 to 16.
Citation Information
Patent Citations
Promoting voice actions to hotwords
US20190115026A1