On-device monitoring and analysis of on-device machine learning models
By monitoring and analyzing machine learning model performance on the user device, generating and transmitting performance indicators, the problem of difficulty in detecting ML performance degradation on the user device in the prior art is solved, real-time monitoring and analysis of ML model performance on the user device is realized, ensuring the stability and security of ML performance.
Patent Information
- Application Number
- CN202280101503.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-03
- Publication Date
- 2025-06-10
AI Technical Summary
The prior art is difficult to effectively monitor and analyze machine learning model performance on user devices, especially in the local environment of user devices, where ML performance deterioration or deterioration in small subsets of user devices cannot be captured.
By implementing monitoring and analysis of machine learning models on the device on the user device, performance data representing the performance characteristics of machine learning models on the device are obtained, performance metrics are generated, and these metrics are transmitted to the remote system without exposing the input data or predicted output content.
Real-time monitoring and analysis of the performance of machine learning models on user devices is realized, and ML regression or performance degradation on the device can be detected, avoiding the deterioration of ML performance to the vast majority of users, and reducing the risk of data exposure to remote systems.
Smart Images

Figure CN120129912A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to machine learning (ML) models on a device. Background Art
[0002] The use of machine learning (ML) is becoming increasingly common. These ML models can be configured and trained to generate any one of various predictions, estimations, classifications, identifications, etc. based on input data. For example, predicting what a user said (i.e., transcribing) based on audio data that captures the user's spoken words. In other examples, ML models are used to identify objects or people in an image, identify media content, and analyze medical images. Summary of the Invention
[0003] One aspect of the present disclosure provides a method that includes obtaining a pre-trained machine learning model from a remote system, receiving input data captured by a user device, and processing using a machine learning model on the device corresponding to the pre-trained machine learning model to generate a plurality of predicted outputs. The method further includes obtaining performance data representing one or more performance characteristics of the machine learning model on the device, where the one or more performance characteristics characterize the performance of the machine learning model on the device based on the plurality of predicted outputs. The method further includes: generating one or more performance metrics for the machine learning model on the device using the performance data without exposing the content of the input data or the plurality of predicted outputs to the remote system, and transmitting the one or more performance metrics to the remote system.
[0004] Implementations of the present disclosure may include one or more of the following optional features. In some examples, the performance data includes the difference between the plurality of predicted outputs and one or more user corrections to the plurality of predicted outputs. The difference may include the number of edits to the plurality of predicted outputs based on one or more user corrections, and generating one or more performance metrics includes determining an edit rate based on the number of edits. For each specific predicted output among the plurality of predicted outputs, the difference may further include an indication of whether the user corrected the specific predicted output, and generating one or more performance metrics includes determining the incidence of user corrections based on the indication.
[0005] In some implementations, the performance data includes prediction probabilities determined by the machine learning model on the device when generating the plurality of predicted outputs. In some examples, the performance data includes at least one of the following: the amount of time to generate the plurality of predicted outputs, the memory usage to generate the plurality of predicted outputs, or a failure of the machine learning system to execute the machine learning model on the device. In some implementations, obtaining the performance data includes obtaining performance data for a plurality of time steps.
[0006] In some examples, the method further includes updating a machine learning model on a device over time based on multiple predicted outputs and one or more user corrections to the multiple predicted outputs. Performance data may include the prediction accuracy of the machine learning model on the device over time when the user device updates the machine learning model on the device. The performance data may also include the number of parameter values of the machine learning model on the device over time. The prediction accuracy may include an indication that updating the machine learning model on the device has resulted in the machine learning model on the device underlearning or overlearning the user corrections.
[0007] In some implementations, the method further includes: storing a snapshot of the machine learning model on the device when the user device updates the machine learning model on the device, and restoring the machine learning model on the device to the stored snapshot based on one or more of the performance metrics. In some examples, obtaining the machine learning model on the device includes obtaining a specific metric definition for each particular performance metric of one or more performance metrics, the specific metric definition including an indication of the performance data to be obtained related to the particular performance metric, and the logic for generating the particular performance metric. The specific metric definition may also include logic for taking an action based on the value of the particular performance metric. Additionally or alternatively, the logic for taking an action may include at least one of the following: transmitting one or more performance metrics to a remote system, restoring the machine learning model on the device to a previous state, disabling the machine learning mode, stopping the update of the machine learning model on the device, or replacing the machine learning model on the device with a different machine learning model on the device. In some examples, the specific metric definition is generated by a developer who: generated a pre-trained machine learning model, deployed the pre-trained machine learning model to the user device and one or more other user devices via a remote system, received one or more performance metrics from the user device and one or more other user devices via the remote system, and analyzed the one or more performance metrics from the user device and one or more other user devices to evaluate the operation of the pre-trained machine learning model.
[0008] In some implementations, the method further includes: storing one or more performance metrics on the user device, and transmitting the one or more performance metrics to a remote system based on at least one of the following: a periodic schedule, a received request, the value of a particular performance metric of the one or more performance metrics, or an error condition.
[0009] Another aspect of the present disclosure provides a system that includes data processing hardware; and memory hardware communicatively coupled to the data processing hardware, the memory hardware storing instructions that, when executed on the data processing hardware, cause the data processing hardware to perform operations of a method that includes obtaining a pre-trained machine learning model from a remote system, receiving input data captured by the system, and processing the input data using an on-device machine learning model corresponding to the pre-trained machine learning model to generate a plurality of predicted outputs. The method further includes obtaining performance data representing one or more performance characteristics of the on-device machine learning model, the one or more performance characteristics characterizing the performance of the on-device machine learning model based on the plurality of predicted outputs. The method further includes: generating one or more performance metrics for the on-device machine learning model using the performance data without exposing the content of the input data or the plurality of predicted outputs to the remote system, and transmitting the one or more performance metrics to the remote system.
[0010] Implementations of the present disclosure may include one or more of the following optional features. In some examples, the performance data includes the differences between the plurality of predicted outputs and one or more user corrections to the plurality of predicted outputs. The differences may include the number of edits to the plurality of predicted outputs based on one or more user corrections, and generating one or more performance metrics includes determining an edit rate based on the number of edits. For each particular predicted output of the plurality of predicted outputs, the differences may also include an indication of whether the user corrected the particular predicted output, and generating one or more performance metrics includes determining the incidence of user corrections based on the indication.
[0011] In some implementations, the performance data includes the prediction likelihoods determined by the on-device machine learning model when generating the plurality of predicted outputs. In some examples, the performance data includes at least one of the following: the amount of time to generate the plurality of predicted outputs, the memory usage to generate the plurality of predicted outputs, or a failure of the machine learning system to execute the on-device machine learning model. In some implementations, obtaining the performance data includes obtaining the performance data for a plurality of time steps.
[0012] In some examples, the method further includes updating the on-device machine learning model over time based on the plurality of predicted outputs and one or more user corrections to the plurality of predicted outputs. The performance data may include the prediction accuracy of the on-device machine learning model over time when the system updates the on-device machine learning model. The performance data may also include the number of parameter values of the on-device machine learning model over time. The prediction accuracy may include an indication that the update of the on-device machine learning model results in the on-device machine learning model underlearning or overlearning user corrections.
[0013] In some implementations, the method further includes: storing a snapshot of the machine learning model on the storage device when the machine learning model is updated on the system update device, and restoring the machine learning model on the device to the stored snapshot based on one or more of the performance metrics. In some examples, obtaining the machine learning model on the device includes obtaining a specific metric definition for each specific performance metric among one or more performance metrics, the specific metric definition including an indication of the performance data to be obtained related to the specific performance metric, and the logic for generating the specific performance metric. The specific metric definition may further include logic for taking an action based on the value of the specific performance metric. Additionally or alternatively, the logic for taking an action may include at least one of the following: transmitting one or more performance metrics to a remote system, restoring the machine learning model on the device to a previous state, disabling the machine learning mode, stopping the update of the machine learning model on the device, or replacing the machine learning model on the device with a different machine learning model on the device. In some examples, the specific metric definition is generated by a developer who: generated the pre-trained machine learning model, deployed the pre-trained machine learning model to the system and one or more other user devices via a remote system, received one or more performance metrics from the system and one or more other user devices via the remote system, and analyzed one or more performance metrics from the system and one or more other user devices to evaluate the operation of the pre-trained machine learning model.
[0014] In some implementations, the method further includes: storing one or more performance metrics on the user device, and transmitting one or more performance metrics to a remote system based on at least one of the following: a periodic schedule, a received request, the value of a specific performance metric among one or more performance metrics, or an error condition.
[0015] Details of one or more implementations of the present disclosure are set forth in the accompanying drawings and the description below. Other aspects, features, and advantages will be apparent from the description and drawings, and from the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 An example ML system is depicted that utilizes on-device monitoring and analysis of a machine learning (ML) model on a device.
[0017] Figure 2 is Figure 1 a schematic diagram of an example of an on-device ML system of
[0018] Figure 3 is Figure 1 a schematic diagram of an example of an on-device ML monitoring process of
[0019] Figure 4A flowchart of an example arrangement of operations of a computer-implemented method for on-device monitoring and analysis of an ML model on a device.
[0020] Figure 5 A schematic diagram of an example computing device that can be used to implement the systems and methods described herein.
[0021] In the various figures, like reference numerals indicate like elements. Detailed Description
[0022] Conventionally, performance monitoring and analysis of machine learning (ML) models are performed using federated analysis on a central server, which aggregates data related to the use of the ML model reported by a distributed pool of users or client devices (i.e., associated with end-users). While the server can use the aggregated data to measure the impact of ML features (e.g., new or updated ML models) on the entire group of user devices, the use of aggregated data cannot capture regressions (e.g., ML performance degradation or deterioration) in a small subset of user devices. This can be particularly a problem when user devices are configured to personalize a copy of the ML model of the user device based on local data captured by the user device. For example, a user device associated with a particular user can personalize a local copy of an automatic speech recognition (ASR) model to better recognize the particular user's spoken words. However, in some cases, ML model personalization by the user device can result in an ML regression, which, if undetected, can degrade the user's experience even though the ML model personalization improves the ML performance of the vast majority of user devices. A server using data aggregation may also not be able to detect when an ML feature has not been adequately trained (i.e., a model trained using training examples that do not adequately represent the local data of a particular user device). For example, an ASR model trained only using words spoken in English without any accent may perform poorly when attempting to recognize the words of a person with a strong foreign accent in English, even though the ASR model works very well for the vast majority of users. Regressions can be caused by a variety of reasons, such as insufficient learning, overlearning, poor system state, device capabilities, etc. When an ML feature works very well for the vast majority of users, monitoring aggregated metrics cannot detect regressions or poor performance on a few specific devices or for a few specific users. Additionally, even if an ML regression on a few devices or for a few users can be detected from the aggregated data, it may wrongly cause developers to turn off the ML feature or revert to a previous ML model, which may unnecessarily result in a deterioration of the ML performance for the vast majority of users. Moreover, in some cases, the regression can be due to aspects of the device that are unrelated to the ML feature (e.g., device state, processor capabilities, available memory, etc.), such that adjusting the ML feature for other users in response to such cases may be undesirable and may result in an unnecessary deterioration of the ML performance for the vast majority of users.
[0023] Figure 1Is a schematic diagram of an example machine learning (ML) system 100 configured to evaluate and / or debug ML performance using on-device monitoring and analysis of on-device ML models. In the example shown, system 100 includes a remote system 150 and a plurality of client or user devices 110, 110a-n (collectively referred to herein as user devices 110), each device being associated with a corresponding end user 130. The remote system 150 and the user devices 110 are communicatively coupled via one or more communication networks 170 (e.g., any combination of wired and / or wireless local area networks (LANs), wide area networks (WANs), cellular networks, and / or any other type of network).
[0024] Each user device 110 includes an on-device ML system 200 and an on-device ML monitoring process 300. The on-device ML system 200 executes one or more on-device ML models 210, 210O, which are configured to process input data 221 captured by the user device 110 on the user device 110 ( Figure 2 ) to generate a predicted output 222 ( Figure 2 ). Each on-device ML model 210O may include a corresponding one of one or more pre-trained ML models 210, 210T deployed to the user device 110 by the remote system 150. In some examples, the on-device ML system 200 executes the pre-trained ML model 210T and updates or personalizes the pre-trained ML model 210T to generate a corresponding on-device ML model 210O, which may be different from the initial version of the pre-trained ML model 210T deployed to the user device 110 by the remote system 150. That is, as a result of on-device personalization and update of the initial pre-trained ML model 210T deployed to multiple user devices 110, one on-device ML model 210O corresponding to the on-device copy of the deployed ML model 210T of the first user device may be different from another on-device ML model 210O corresponding to the on-device copy of the same deployed ML model 210T of the second user device. In some examples, when the personalized on-device ML model 210O is trained, updated, or otherwise personalized, the on-device ML system 200 saves snapshots of the personalized on-device ML model 210O over time.
[0025] The on-device ML monitoring process 300 is configured to monitor and analyze the ML performance of one or more on-device ML models 210O implemented by the on-device ML system 200. Specifically, the on-device ML monitoring process 300 obtains performance data 250 representing one or more performance characteristics of one or more on-device ML models 210O (e.g., from the on-device ML system 200) ( Figure 2) Calculate one or more user device-specific ML model performance metrics 302 (i.e., on-device ML performance metrics 302) based on performance data, and analyze the on-device ML performance metrics 302 to identify trends in ML performance over time. The on-device ML monitoring process 300 can aggregate the on-device calculated ML performance metrics 302 over time on the user device 110 to populate a database (e.g., securely stored in the data repository 310 of the user device 110) that tracks how a particular ML feature (e.g., ML model) performs on the user device 110. Example performance data 250 includes, but is not limited to: the difference between the predicted output of the on-device ML model and the user's correction of it for multiple time steps or within multiple time steps; the number of edits (e.g., word additions, deletions, replacements in a transcription of spoken words); an indication of whether the predicted output was corrected and / or which predicted outputs were corrected; the prediction likelihood associated with the prediction hypothesis determined by the on-device ML model when generating the predicted output; the processing time to generate the predicted output; the memory usage to generate the predicted output; fault conditions; machine learning system failure conditions; prediction accuracy; the number of parameter values of the ML model that change over time; user indications of whether the correction is overlearning or underlearning (e.g., the user continues to make the same correction or reverts to a previously trained correction); and user feedback. Example ML model performance metrics 302 include, but are not limited to: edit rate (e.g., the frequency and number of edits made to a transcription of spoken words over time, such as word error rate (WER)); the incidence of user corrections; whether the prediction confidence is increasing or decreasing; whether the parameter values of the ML model are jittering; the processor usage trend; and the memory usage trend.
[0026] In some examples, the on-device ML monitoring process 300 is configured to take one or more actions in response to the on-device ML performance metrics 302 or in response to trends based on the on-device ML performance metrics 302. That is, the on-device ML monitoring process 300 can locally adjust the operations performed by the on-device ML system 200 or the on-device ML model implemented by the on-device ML system 200 (i.e., measure locally and take action locally) in response to the local on-device ML performance metrics 302. In some implementations, the on-device ML monitoring process 300 uses the on-device ML model performance trends over time to determine whether to: turn on or off the on-device ML function, disable the ML function, reset the state of the on-device ML model 210O, restore to a previous on-device ML model snapshot 210, 210S ( Figure 2)(e.g., restoring to the previous version of the ML model on the device with the best performance), stopping the update of the ML model 210O on the device, replacing the ML model 210O on the device with a different ML model, etc. For example, in the case of on-device personalization of the ML model 210O corresponding to the ASR model, the on-device ML monitoring process 300 analyzes and tracks the speech recognition performance of the on-device personalized ASR model (e.g., measured by the number of transcription corrections made by the user) over a period of time, i.e., to determine the WER of the ASR model. Thus, when the on-device ML monitoring process 300 detects a performance regression of the ASR model (e.g., a deterioration in speech recognition accuracy is associated with an increase in WER), the on-device ML monitoring process 300 can disable future personalization, restore to the previously trained ASR model, restore to the base ASR model (e.g., non-personalized model), submit a defect report, etc. The on-device ML monitoring process 300 can also respond to user input. For example, the user provides an indication that any on-device ML model updates made in the past N days should be discarded because any user corrections provided during these days were made by a party other than the user 130 associated with the user device 110 (e.g., a child got hold of the parent's user device 110).
[0027] The on-device ML monitoring process 300 can also provide the on-device ML monitoring and analysis results (e.g., the on-device ML performance metrics 302) to the remote system 150 in, for example, a periodic report, a response to a query, a debug log, or a defect report, which ML developers can use to identify and debug ML model problems that affect even a small subset of the user devices 110 only. In some examples, the on-device ML monitoring process 300 calculates the on-device ML performance metrics 302 such that the on-device ML performance metrics 302 do not contain or disclose any content of the captured input data or the predicted output (i.e., anonymize and / or sanitize the performance metrics 302).
[0028] In some implementations, the ML developer of the deployed ML model obtains a pre-trained machine learning model from a remote system, receives input data captured by the user device, and processes it using an on-device machine learning model corresponding to the pre-trained machine learning model to generate multiple predicted outputs. The method further includes obtaining performance data representing one or more performance characteristics of the on-device machine learning model, where the one or more performance characteristics characterize the performance of the on-device machine learning model based on the multiple predicted outputs. The method further includes: generating one or more performance metrics for the on-device machine learning model using the performance data without exposing the content of the input data or the multiple predicted outputs to the remote system, and transmitting the one or more performance metrics to the remote system.
[0029] Implementations of the present disclosure may include one or more of the following optional features. In some examples, the performance data includes the difference between multiple predicted outputs and one or more user corrections to the multiple predicted outputs. The difference may include the number of edits to the multiple predicted outputs based on one or more user corrections, and generating one or more performance metrics includes determining an edit rate based on the number of edits. For each particular predicted output of the multiple predicted outputs, the difference may also include an indication of whether the user corrected the particular predicted output, and generating one or more performance metrics includes determining the incidence of user corrections based on the indication.
[0030] In some implementations, the performance data includes the prediction likelihoods determined by the machine learning model on the device when generating the multiple predicted outputs. In some examples, the performance data includes at least one of the following: the amount of time used to generate the multiple predicted outputs, the memory usage to generate the multiple predicted outputs, or the failure of the machine learning system to execute the machine learning model on the device. In some implementations, obtaining the performance data includes obtaining the performance data for multiple time steps.
[0031] In some examples, the method further includes updating the machine learning model on the device over time based on the multiple predicted outputs and one or more user corrections to the multiple predicted outputs. The performance data may include the prediction accuracy of the machine learning model on the device over time when the user device updates the machine learning model on the device. The performance data may also include the number of parameter values of the machine learning model on the device over time. The prediction accuracy may include an indication that the update of the machine learning model on the device results in the machine learning model on the device underlearning or overlearning the user corrections.
[0032] In some implementations, the method further includes: storing a snapshot of the machine learning model on the device when the user device updates the machine learning model on the device, and restoring the machine learning model on the device to the stored snapshot based on one or more of the performance metrics. In some examples, obtaining the machine learning model on the device includes obtaining a specific metric definition for each specific performance metric of one or more performance metrics, the specific metric definition including an indication of the performance data to be obtained that is related to the specific performance metric, and the logic for generating the specific performance metric. The specific metric definition may further include the logic for taking an action based on the value of the specific performance metric. Additionally or alternatively, the logic for taking an action may include at least one of the following: transmitting one or more performance metrics to a remote system, restoring the machine learning model on the device to a previous state, disabling the machine learning mode, stopping the update of the machine learning model on the device, or replacing the machine learning model on the device with a different machine learning model on the device. In some examples, the specific metric definition is generated by a developer who: generated the pre-trained machine learning model, deployed the pre-trained machine learning model to the user device and one or more other user devices via a remote system, received one or more performance metrics from the user device and one or more other user devices via the remote system, and analyzed one or more performance metrics from the user device and one or more other user devices to evaluate the operation of the pre-trained machine learning model.
[0033] In some implementations, the method further includes: storing one or more performance metrics on the user device, and transmitting the one or more performance metrics to a remote system based on at least one of: a periodic schedule, a received request, a value of a particular performance metric among the one or more performance metrics, or an error condition. Providing one or more metric definitions 304 along with the deployed ML model 210T, the one or more metric definitions defining for the on-device ML monitoring process 300 on the device: how to calculate, track, and report the on-device ML performance metrics 302 and which on-device ML performance metrics 302 to calculate, track, and report, and automated actions to take in response to the on-device ML performance metrics 302 or trends in the on-device ML performance metrics 302. That is, the ML developer can a priori define a preemption mechanism, a recovery mechanism, or an escape mechanism in case the deployed ML model 210O does not execute on a particular user device 110 in the manner expected by the ML developer (e.g., fails to meet one or more performance thresholds). In this way, the ML developer can configure the on-device ML monitoring process 300 of the user device 110 to reduce the likelihood that the deployed ML features (e.g., the ML model 210T) negatively impact a small subset of the user devices 110 and obtain the on-device ML performance metrics 302 such that even if the ML performance issue only affects a small subset of the user devices 110, the ML developer can identify the root cause of the problem. In some implementations, the on-device ML monitoring process 300 equips or configures the on-device ML system 200 in response to the metric definitions 304 to capture the performance data required to calculate a particular ML performance metric 302 and store the captured performance data on the user device (e.g., securely store it in the data repository 310). In some examples, when the user device 110 is idle (e.g., the user is not currently interacting with the user device 110 or the current resource utilization of the user device 110 meets a threshold), the on-device ML monitoring process 300 processes the captured performance data to calculate the on-device ML performance metrics 302 so as not to interfere with other functions of the user device 110.
[0034] As used herein, "on-device" means that a particular user device 110 performs / conducts a process or function on behalf of an end user 130 of the user device 110 completely independently of computing and storage resources implemented on a remote system 150 or any other user device 110. For example, an on-device ML model 210O refers to an ML model implemented by or on the user device 110. This is in contrast to sending input data captured by the user device 110 to a central server (e.g., the remote system 150), which executes an ML model on behalf of the user device 110 and one or more other user devices 110, computes a predicted output for the input data using the ML model, and returns the predicted output to the user device 110. Similarly, on-device monitoring and analysis of the on-device ML model 210O refers to the monitoring and analysis of the on-device ML model 210O executed locally by or on the user device 110. That is, the user device 110 performs the monitoring and analysis of the on-device ML model 210O without sending any performance-related data collected by the user device 110 to the remote system 150. In some implementations, data processing hardware 112 (e.g., a programmable processor) of the user device 110 executes instructions stored on a memory hardware 114 of the user device 110 to implement the on-device ML system 200 and the on-device ML monitoring process 300. Additionally or alternatively, the data processing hardware 112 may implement dedicated data hardware (e.g., a tensor processing unit (TPU)) for executing the on-device ML model 210O on the user device 110.
[0035] The user device 110 may correspond to any personal computing device associated with a user and capable of receiving input, performing processing, and providing output. Some examples of the user device 110 include, but are not limited to, mobile devices (e.g., mobile phones, tablet computers, laptop computers, etc.), computers, wearable devices (e.g., smart watches), smart home appliances, Internet of Things (IoT) devices, in-vehicle infotainment systems, smart displays, smart speakers, etc. Each user device 110 includes data processing hardware 112 and memory hardware 114 that communicates with the data processing hardware 112. The memory hardware 114 stores instructions that, when executed by the data processing hardware 112, cause the data processing hardware 112 or more generally the user device 110 to perform one or more operations. Each user device 110 includes or is coupled to one or more input systems 116 (e.g., audio capture devices such as microphone 116a, virtual keywords, keyboards, etc.) to capture, record, receive, or otherwise obtain user input (e.g., spoken words) for the user device 110. Each user device 110 also includes or is coupled to one or more output systems 118 (e.g., speaker 118a, screen 118b, etc.) to output or otherwise provide the output of the user device 110 (e.g., the predicted output of the on-device ASR model) to the user 130. The input system 116 may also be used to obtain user input from other users 130, devices, systems, etc. The output system 118 may also be used to provide output to other users 130, devices, systems, etc.
[0036] In one example, the on-device ML system 200 of the example user device 110a implements an on-device ML model 210O that includes an ASR model (not shown for clarity). The ASR model can include a recurrent neural network transducer (RNN-T) model and an optional rescoring component, each residing on the user device 110a. The input system 116 of the example user device 110a includes an audio subsystem configured to receive utterance 135 spoken by the user 130a and captured by an audio capture device 116a (e.g., one or more microphones), and convert the captured utterance 135 into a corresponding digital format associated with input audio data that can be digitally input into and processed by the ASR system. Thereafter, the RNN-T model receives the audio data corresponding to the utterance 135 as input and generates / predicts a corresponding transcription 136 (e.g., recognition result / hypothesis) of the utterance 135 as output. The RNN-T model can perform streaming speech recognition to produce initial speech recognition results 136, 136a, and the rescoring component can update the initial speech recognition result 136a (i.e., rescore it) to produce a final speech recognition result 136, 136b. Thereafter, when, for example, natural language processing (NLP) performed on the final speech recognition result 136b identifies that the spoken utterance 135 is, for example, a command or a query, the user device 110a can provide the final speech recognition result 136b to a downstream application (e.g., digital assistant 126) to execute the command or identify a response to the query (e.g., response 138).
[0037] In another example, the on-device ML system 200 of the example user device 110 implements a text-to-speech (TTS) model 210O configured to convert input text into a synthetic speech representation that can be used by a vocoder / synthesizer (not shown for illustration clarity) to output, in an audible manner, synthetic speech corresponding to the input text. The input system 116 of the example user device 110 includes a text input subsystem (e.g., a keyboard or a virtual keyboard) configured to receive input text (i.e., input data representing one or more characters and / or words) from the user 130. Alternatively, the input text is received from the digital assistant 126 with which the user 130 is interacting. For example, the user 130 types the input text and the digital assistant 126 responds with a synthetic speech output. In some examples, the on-device ML system 200 additionally implements an ASR model, such as the ASR described above, such that the user 130b interacts with the digital assistant 126 via spoken utterance input and synthetic speech output.
[0038] The on-device ML system 200 can similarly implement any number and variety of other on-device ML models 210O. For example, the on-device ML system 200 implements, but is not limited to, an image recognition ML model, a classification ML model, a medical diagnosis ML model, an object recognition ML model, a person recognition ML model, a speaker recognition ML model, a media content recognition ML model, a speech-to-speech model, a language model, a language translation model, a machine translation model, or any other type of ML model trained via ML to generate a predictive output based on a received input.
[0039] Reference Figure 1 , the remote system 150 includes data processing hardware 152 and memory hardware 154 communicatively coupled to the data processing hardware 152. The memory hardware 154 stores instructions that, when executed by the data processing hardware 152, cause the data processing hardware 152 to perform one or more operations, such as those disclosed herein. In some examples, the remote system 150 is provided by an ML model developer. Alternatively, the remote system 150 is a central server that trains and deploys the ML model 210T and the metric definition 304, and obtains on-device ML performance metrics 302 from user devices 110 that execute the deployed ML model 210T on behalf of multiple different ML model developers.
[0040] An example remote system 150 includes an ML model data repository 156 for storing the ML model 210T deployed by the remote system 150 to the user device 110. In some examples, the metric definition 304 is stored in the ML model data repository 156 together with its corresponding ML model 210T.
[0041] In the example shown, the remote system 150 includes a metric aggregation process 157 that receives or obtains on-device ML performance metrics 302 from the user device 110 and stores the on-device ML performance metrics 302 in a metric data repository 158. In some implementations, the metric aggregation process 157 populates a database stored on the data repository 158 to track over time how a particular ML feature performs on various user devices 110. In some examples, the metric aggregation process 157 aggregates the on-device ML performance metrics 302 received from various user devices 110 to determine the performance of the ML features for the entire population of user devices 110. Additionally or alternatively, the metric aggregation process 157 uses the on-device ML performance metrics 302 of a particular user device 110 to track the performance of a particular ML feature on that particular user device 110. In some examples, when the ML performance metrics 302 of a particular user device 110 degrade, the metric aggregation process 157 receives the on-device ML performance metrics 302 from the particular user device 110 via a debug log or a defect report, such that the metric aggregation process 157 becomes aware that the on-device ML model implemented by the particular user device 110 is not performing as expected by the ML developer of the ML model.
[0042] The remote system 150 includes an application programming interface (API) 159 or other user interface for enabling an ML model developer to provide an ML model 210T to be deployed by the remote system 150 to the user device 110, as well as one or more corresponding metric definitions 304 for the ML model 210T that will also be provided to the user device 110. In some examples, the remote system 150 stores a database of the metric definitions 304 such that the ML model developer can select only which of the metric definitions 304 in the database are to be used with a particular deployed ML model 210T. Additionally or alternatively, the user device 110 stores a database of the metric definitions 304 such that the ML model developer can identify only which on-device ML performance metrics 304 or their trends are to be calculated and tracked for the user device 110.
[0043] Figure 2 is Figure 1 A schematic diagram of an example of an on-device ML system 200. The on-device ML system 200 includes an ML model data repository 210 for storing one or more deployed ML models 210, 210Ta-Tn received from the remote system 150, one or more on-device ML models 210, 210Oa-On, and / or one or more ML model snapshots 210, 210Sa-Sn (i.e., snapshots of the on-device ML model 210O). The on-device ML model 210O may include a personalized on-device ML model, i.e., a personalized copy or version of the deployed ML model 210T.
[0044] The on-device ML system 200 includes one or more on-device ML engines 220, 220a-n, which are configured to execute an on-device ML model 210O to process input data 221 captured by the input system 116 to generate a predicted output 222, which can be output, for example, by the output system 118 or provided for use by the user device 110 or the digital assistant 126 when performing downstream operations. In some implementations, the data processing hardware 112 (e.g., a programmable processor) of the user device 110 executes instructions stored on the memory hardware 114 of the user device 110 to implement one or more of the on-device ML engines 220 for executing the on-device ML model 210O. In some examples, the data processing hardware 112 includes dedicated data hardware (e.g., a tensor processing unit (TPU)) to implement the on-device ML engine 220 for executing the on-device ML model 210O. In some examples, the on-device ML engine 220 executes more than one on-device ML model 210O.
[0045] The on-device ML system 200 includes a model selection process 230 that selects an on-device ML model for execution by the on-device ML engine 220 in response to an input 232 received from the on-device ML monitoring process 300. For example, the on-device ML monitoring process 300 selects the current version of the on-device ML model 210O or a previous snapshot 210S of the on-device ML model 210O for execution by the on-device ML engine 220. The on-device ML monitoring process 300 can also control the model selection process 230 to disable the on-device ML model 210O or snapshot 210S such that the on-device ML engine 220 no longer executes the disabled on-device ML model 210O or snapshot 210S.
[0046] In some examples, the on-device ML system 200 includes an on-device ML training engine 240 that personalizes or updates the on-device ML model 210O based on, for example, captured input data 221, predicted output 222, prediction-related data 242 from the on-device ML engine 220 (e.g., prediction hypotheses, prediction likelihoods, etc.), and / or user input 244 (e.g., user corrections).
[0047] In the example shown, the on-device ML system 200 is equipped (e.g., configured) to capture and provide on-device ML performance data 250 to the on-device ML monitoring process 300, which represents one or more performance characteristics of one or more on-device ML models 210O executed on the user device 110. In some implementations, the on-device ML monitoring process 300 configures the on-device ML system 200 to capture and report to the on-device ML monitoring system 300 the performance data 250 for each on-device ML model 210. In some additional implementations, the deployed ML model 210T includes a definition of the performance data 250 to be captured. In other implementations, the on-device ML system 200 is configured to collect and provide to the on-device ML monitoring process 300 a set of default or normalized performance data 250 for each executed on-device ML model 210O, and the on-device ML monitoring process 300 determines which performance data 250 to store and use to calculate the ML performance metric 302.
[0048] Example performance data includes, but is not limited to: the difference between the predicted output of the on-device ML model and the user's correction thereof over multiple time steps or within multiple time steps; the number of edits (e.g., word additions, word deletions, word replacements to a transcription of spoken words); an indication of whether the predicted output was corrected and / or which predicted outputs were corrected; the prediction likelihood associated with the prediction hypothesis determined by the on-device ML model in generating the predicted output; the processing time used to generate the predicted output; the memory usage used to generate the predicted output; fault conditions; machine learning system failure conditions; prediction accuracy; the number of parameter values of the ML model that change over time; a user indication of whether the correction is overfitting or underfitting (e.g., the user continues to make the same correction or reverts to a previously trained correction); and user feedback.
[0049] Although Figure 2 an example on-device ML system 200 is shown, Figure 2 one or more of the elements and processes shown therein may be combined, divided, rearranged, omitted, eliminated, or implemented in any other way. Additionally, the on-device ML system 200 may include one or more elements or processes that supplement or replace the elements or processes shown in Figure 2 or may include more than one of any or all of the elements and processes shown therein.
[0050] Figure 3It is a schematic diagram of an example of the on-device ML monitoring process 300. The on-device ML monitoring process 300 includes a data collection process 320 that is used to obtain or receive ML performance data 250 from the on-device ML system 200 over time and at multiple time steps, and store the ML performance data 250 in the metric data repository 310. The data repository 310 can store or retain the performance data 250 for any period of time. For example, during the duration of a reporting period - a duration defined by the ML model developer - until the data repository 310 is full and older performance data 250 is discarded, until the user device 110 is restarted, etc. In some implementations, the data collection process 320 configures the on-device ML system 200 to provide specific performance data 250 for a specific on-device ML model 210O. In some additional implementations, the performance data 250 includes a set of default or normalized performance data 250, and the data collection process 320 selects which of the performance data 250 to store in the data repository 310.
[0051] The on-device monitoring process 300 includes an analysis configuration process 330 that, in response to a metric definition 304, configures the data collection process 320 and / or the on-device ML system 200 to collect specific ML performance data 250 and store the performance data 250 in the data repository 310. The analysis configuration process 330 also configures a set of one or more metric logics 342 for calculating and reporting ML performance metrics 302 and / or their trends in response to the metric definition 304. The set of metric logics 342 can also list actions to be taken in response to the ML performance metrics 302 and / or their trends.
[0052] The on-device monitoring process 300 includes a metric monitoring and reporting process 340 that is used to execute the set of metric logics 342 to calculate and track ML performance metrics 302 and / or their trends, and / or take actions in response to the ML performance metrics 302 and / or their trends.
[0053] For example, one of the metric logics in the metric logic set 342 can define how the metric monitoring and reporting process 340 calculates one or more specific ML performance metrics 302. Specifically, the metric logic 342 can define: what performance data 250 the metric monitoring and reporting process 340 is to use, what logic or equations the metric monitoring and reporting process 340 is to use to process the performance data 250 to calculate the specific ML performance metric 302, and / or how the metric monitoring and reporting process 340 aggregates the specific ML performance metric 302 over time to identify and track trends in the ML performance metric 302. Example ML model performance metrics 302 include, but are not limited to, edit rate (e.g., the frequency and number of edits made to transcriptions of spoken utterances over time, such as word error rate (WER)); the incidence of user corrections; whether the prediction confidence is increasing or decreasing; whether the parameter values of the ML model are jittering; processor usage trends; and memory usage trends. For example, in the case of on-device personalization of an ASR model on a device, one of the metric logics in the metric logic set 342 can cause the metric monitoring and reporting process 340 to analyze and track the speech recognition performance of the on-device personalized ASR model (e.g., as measured by the number of transcription corrections made by the user) over a period of time. In some examples, the metric monitoring and reporting process 340 calculates the on-device ML performance metric 302 such that the on-device ML performance metric 302 does not contain or (e.g., to the remote system 150) disclose any content of the captured input data or predicted output.
[0054] One of the metric logics in the metric logic set 342 can additionally or alternatively define when and / or how the metric monitoring and reporting process 340 stores, records, and / or reports the specific ML performance metric 302. For example, the metric logic set 342 defines that the ML performance metric 302 will be provided via periodic reports, responses to queries, debug logs, and / or defect reports.
[0055] One metric logic in the set of metric logics 342 can additionally or alternatively define one or more specific actions taken by the metric monitoring and reporting process 340, and the logic used by the metric monitoring and reporting process 340 to determine when to take that specific action. Example specific operations include, but are not limited to: turning on or off the ML function on the device, disabling the ML function, resetting the state of the ML model on the device, restoring the ML model on the device to a previous snapshot of the ML model on the device (e.g., restoring to the previous version with the best performance of the ML model on the device), stopping the update of the ML model on the device, and replacing the ML model on the device with a different ML model. For example, when the metric monitoring and reporting process 340 detects a performance regression of the ASR model (e.g., deterioration of speech recognition accuracy), the set of metric logics 342 causes the metric monitoring and reporting process 340 to disable future personalization, restore to a previously trained ASR model, restore to the base ASR model (non-personalized model), submit a defect report, etc.
[0056] One metric logic in the set of metric logics 342 can additionally and / or alternatively define how the metric monitoring and reporting process 340 responds to user input. For example, the user can provide an indication that any ML model updates on the device performed in the past N days should be discarded because any user corrections provided during those days were made by someone other than the user 130 associated with the user device 110 (e.g., a child got hold of the parent's user device 110), such that the set of metric logics 342 causes the metric monitoring and reporting process 340 to restore the ML model on the device to a previous version.
[0057] Although Figure 3 the example device - based ML monitoring process 300 is shown, Figure 3 one or more of the elements and processes shown in Figure 3 can be combined, divided, rearranged, omitted, eliminated, or implemented in any other way. Additionally, the device - based ML monitoring process 300 can include one or more elements or processes that supplement or replace the elements or processes shown in
[0058] Figure 4It is a flowchart of an exemplary arrangement of operations of a computer-implemented method 400 executed by a user device 110 for on-device monitoring and analysis of an ML model 210O on a device. At operation 402, method 400 includes obtaining a pre-trained machine learning model 210T from a remote system 150. At operation 404, method 400 includes receiving input data 221 captured by user device 110. At operation 406, method 400 includes processing using an on-device ML model 210O corresponding to the pre-trained ML model 210T to generate a plurality of predicted outputs 222.
[0059] At operation 408, method 400 includes obtaining performance data 250 representing one or more performance characteristics of the on-device ML model 210O, the one or more performance characteristics characterizing the performance of the on-device ML model 210O based on the plurality of predicted outputs 222. At operation 410, method 400 includes using the performance data 250 to generate one or more performance metrics 302 of the on-device ML model 210O without exposing the content of the input data 221 or the plurality of predicted outputs 222 to the remote system 150. Method 400 includes transmitting one or more performance metrics 302 to the remote system 150 at operation 412.
[0060] Figure 5 It is a schematic diagram of an exemplary computing device 500 that can be used to implement the systems and methods described in this document. Computing device 500 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. The components shown here, their connections and relationships, and their functions are intended to be exemplary only and are not intended to limit the implementation of the inventions described and / or claimed in this document.
[0061] The computing device 500 includes a processor 510 (i.e., data processing hardware) that can be used to implement data processing hardware 112 and / or 152, a memory 520 (i.e., memory hardware) that can be used to implement memory hardware 114 and / or 154, a storage device 530 (i.e., memory hardware) that can be used to implement memory hardware 114 and / or 154, a high-speed interface / controller 540 connected to the memory 520 and the high-speed expansion port 550, and a low-speed interface / controller 560 connected to the low-speed bus 570 and the storage device 530. Each of the components 510, 520, 530, 540, 550, and 560 is interconnected using various buses and can be mounted on a common motherboard or otherwise as appropriate. The processor 510 can process instructions for execution within the computing device 500, including instructions stored in the memory 520 or on the storage device 530, to display graphical information of a graphical user interface (GUI) on an external input / output device (such as a display 580 coupled to the high-speed interface 540). In other implementations, multiple processors and / or multiple buses and multiple memories and multiple types of memories can be used as appropriate. Additionally, multiple computing devices 500 can be connected, where each device provides a portion of the necessary operations (e.g., as a server group, blade server cluster, or multi-processor system).
[0062] The memory 520 stores information non-temporarily within the computing device 500. The memory 520 can be a computer-readable medium, a volatile memory unit, or a non-volatile memory unit. The non-temporary memory 520 can be a physical device for temporarily or permanently storing programs (e.g., sequences of instructions) or data (e.g., program state information) for use by the computing device 500. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electrically erasable programmable read-only memory (EEPROM) (e.g., commonly used for firmware such as a boot program). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase change memory (PCM), and magnetic disks or tapes.
[0063] The storage device 530 can provide large-capacity storage for the computing device 500. In some implementations, the storage device 530 is a computer-readable medium. In various different implementations, the storage device 530 can be a floppy disk device, a hard disk device, an optical disk device, or a magnetic tape device, a flash memory, or other similar solid-state memory devices, or an array of devices (including devices in a storage area network or other configurations). In additional implementations, the computer program product is tangibly embodied in an information carrier. The computer program product contains instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer or machine-readable medium, such as the memory 520, the storage device 530, or the memory on the processor 510.
[0064] The high-speed controller 540 manages the bandwidth-intensive operations of the computing device 500, while the low-speed controller 560 manages the lower bandwidth-intensive operations. Such a division of responsibilities is merely exemplary. In some implementations, the high-speed controller 540 is coupled to the memory 520, the display 580 (e.g., via a graphics processor or accelerator), and a high-speed expansion port 550 that can accept various expansion cards (not shown). In some implementations, the low-speed controller 560 is coupled to the storage device 530 and a low-speed expansion port 590. The low-speed expansion port 590 (which can include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet)) can be coupled to one or more input / output devices (such as a keyboard, a pointing device, a scanner) or a networking device (such as a switch or a router), for example, via a network adapter.
[0065] The computing device 500 can be implemented in many different forms, as shown in the figure. For example, it can be implemented as a standard server 500a or implemented multiple times in a group of such servers 500a, be implemented as a laptop computer 500b, or be implemented as part of a rack server system 500c.
[0066] Various implementations of the systems and techniques described herein can be implemented in digital electronic and / or optical circuitry, integrated circuit systems, specially designed ASICs (application-specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system, the programmable system including at least one programmable processor, which can be dedicated or general-purpose and can be coupled to receive data and instructions from a storage system, at least one input device, and at least one output device and to transmit data and instructions to the storage system, at least one input device, and at least one output device.
[0067] A software application (i.e., software resource) can be computer software that instructs a computing device to perform tasks. In some examples, a software application can be referred to as an "application", "app", or "program". Example applications include, but are not limited to, system diagnostic applications, system management applications, system maintenance applications, word processing applications, spreadsheet applications, messaging applications, media streaming applications, social networking applications, and gaming applications.
[0068] These computer programs (also referred to as programs, software, software applications, or code) include machine instructions for a programmable processor and can be implemented using high-level programming and / or object-oriented programming languages and / or assembly / machine languages. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, non-transitory computer-readable medium, device, and / or apparatus (e.g., a disk, optical disk, memory, programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.
[0069] The processes and logical flows described in this specification can be performed by one or more programmable processors (also referred to as data processing hardware) that execute one or more computer programs to perform functions by operating on input data and generating output. The processes and logical flows can also be performed by dedicated logic circuitry, such as an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit). By way of example, processors suitable for executing computer programs include both general and special purpose microprocessors, and any one or more processors of any kind of digital computer. Generally, a processor will receive instructions and data from a read-only memory or a random access memory or both. The basic elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or operatively coupled to receive data from one or more mass storage devices or to transfer data to one or more mass storage devices, or both. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, by way of example including semiconductor memory devices (e.g., EPROM, EEPROM, and flash memory devices), magnetic disks (e.g., internal hard disks or removable disks), magneto-optical disks, and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, dedicated logic circuitry.
[0070] To provide interaction with a user, one or more aspects of the present disclosure may be implemented on a computer having a display device (e.g., a CRT (cathode ray tube), an LCD (liquid crystal display) monitor, or a touch screen) for displaying information to the user and possibly a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback, such as, for example, visual feedback, auditory feedback, or tactile feedback; and input from the user may be received in any form, including sound, voice, or tactile input. Additionally, the computer may interact with the user by sending documents to and receiving documents from the device used by the user; for example, by sending a web page to a web browser on the user's client device in response to a request received from the web browser.
[0071] Unless there is a clear contrary statement, the phrase "at least one of A, B, or C" is intended to refer to any combination or sub-group of A, B, C, such as: (1) only at least one A; (2) only at least one B; (3) only at least one C; (4) at least one A and at least one B; (5) at least one A and at least one C; (6) at least one B and at least one C; and (7) at least one A and at least one B and at least one C. Further, unless there is a clear contrary statement, the phrase "at least one of A, B, and C" is intended to refer to any combination or sub-group of A, B, C, such as: (1) only at least one A; (2) only at least one B; (3) only at least one C; (4) at least one A and at least one B; (5) at least one A and at least one C; (6) at least one B and at least one C; and (7) at least one A and at least one B and at least one C.
[0072] A variety of implementations have been described. However, it should be understood that various modifications can be made without departing from the spirit and scope of the present disclosure. Accordingly, other implementations are within the scope of the following claims.
Claims
1. A computer-implemented method (400), wherein, when the computer-implemented method is executed on data processing hardware (510) of a user device (110), the data processing hardware (510) is caused to perform operations, the operations including: obtaining a pre-trained machine learning model (210O) from a remote system (150); receiving input data (221) captured by the user device (110); processing the input data (221) using a machine learning model on the device (210O) corresponding to the pre-trained machine learning model (210O) to generate a plurality of predicted outputs (222); obtaining performance data (250) representing one or more performance characteristics of the machine learning model on the device (210O), the one or more performance characteristics characterizing the performance of the machine learning model on the device (210O) based on the plurality of predicted outputs (222); generating one or more performance metrics (302) of the machine learning model on the device (210O) using the performance data (250) without exposing the content of the input data (221) or the plurality of predicted outputs (222) to the remote system (150); and transmitting the one or more performance metrics (302) to the remote system (150).
2. The computer-implemented method (400) according to claim 1, wherein, the performance data (250) includes the difference between the plurality of predicted outputs (222) and one or more user corrections to the plurality of predicted outputs (222).
3. The computer-implemented method (400) according to claim 2, wherein: the difference includes the number of edits to the plurality of predicted outputs (222) based on the one or more user corrections; and generating the one or more performance metrics (302) includes determining an edit rate based on the number of edits.
4. The computer-implemented method (400) according to claim 2 or 3, wherein: for each specific predicted output (222) among the plurality of predicted outputs (222), the difference includes an indication of whether the user has corrected the specific predicted output (222); and generating the one or more performance metrics (302) includes determining the incidence of user corrections based on the indication.
5. The computer-implemented method (400) according to any one of claims 1 to 4, wherein, the performance data (250) includes prediction probabilities determined by the machine learning model on the device (210O) when generating the plurality of predicted outputs (222).
6. The computer-implemented method (400) according to any one of claims 1 to 5, wherein, the performance data (250) includes at least one of the following: the amount of time to generate the plurality of predicted outputs (222); the memory usage to generate the plurality of predicted outputs (222); or The machine learning system executes a failure of the machine learning model (210O) on the device.
7. The computer-implemented method (400) according to any one of claims 1 to 6, wherein, obtaining the performance data (250) includes obtaining the performance data (250) for a plurality of time steps.
8. The computer-implemented method (400) according to any one of claims 1 to 7, wherein: the operation further includes updating the machine learning model (210O) on the device over time based on the plurality of predicted outputs (222) and one or more user corrections to the plurality of predicted outputs (222); and the performance data (250) includes the prediction accuracy of the machine learning model (210O) on the device over time when the user device (110) updates the machine learning model (210O) on the device.
9. The computer-implemented method (400) according to claim 8, wherein, the performance data (250) further includes the number of parameter values of the machine learning model (210O) on the device over time.
10. The computer-implemented method (400) according to claim 8 or 9, wherein, the prediction accuracy includes an indication that the update of the machine learning model (210O) on the device results in the machine learning model (210O) on the device learning insufficiently or overly from the user correction.
11. The computer-implemented method (400) according to any one of claims 1 to 10, wherein, the operation further includes: when the user device updates the machine learning model (210O) on the device, storing a snapshot of the machine learning model (210O) on the device; and restoring the machine learning model (210O) on the device to the stored snapshot based on one or more of the performance metrics (302).
12. The computer-implemented method (400) according to any one of claims 1 to 11, wherein, obtaining the machine learning model (210O) on the device includes obtaining a specific metric definition (304) for each specific performance metric (302) among the one or more performance metrics (302), the specific metric definition including: an indication of the performance data (250) related to the specific performance metric (302) to be obtained; and logic (342) for generating the specific performance metric (302).
13. The computer-implemented method (400) according to claim 12, wherein, the specific metric definition (304) further includes logic (342) for taking an action based on the value of the specific performance metric (302).
14. The computer-implemented method (400) according to claim 13, wherein, the logic (342) for taking the action causes the data processing hardware (510) to perform at least one of the following: Transmit the specific performance metric (302) to the remote system (150); Restore the machine learning model (210O) on the device to a previous state; Disable the machine learning model (210O) on the device; Stop updating the machine learning model (210O) on the device; or Replace the machine learning model (210O) on the device with a different machine learning model (210O) on the device.
15. The computer-implemented method (400) according to any one of claims 12 to 14, wherein, the specific metric definition (304) is generated by a developer who: generated the pre-trained machine learning model (210O); deployed the pre-trained machine learning model (210T) to the user device (110) and one or more other user devices (110) via the remote system (150); received the one or more performance metrics (302) from the user device (110) and the one or more other user devices (110) via the remote system (150); and analyzed the one or more performance metrics (302) from the user device (110) and the one or more other user devices (110) to evaluate the operation of the pre-trained machine learning model (210T).
16. The computer-implemented method (400) according to any one of claims 1 to 15, wherein, the operation further includes: storing the one or more performance metrics (302) on the user device (110); and transmitting the one or more performance metrics (302) to the remote system (150) based on at least one of: a periodic schedule, a received request, a value of a specific performance metric (302) among the one or more performance metrics (302), or an error condition.
17. A system (100), wherein, comprises: data processing hardware (510); and memory hardware that communicates with the data processing hardware (510), the memory hardware storing instructions that, when executed on the data processing hardware (510), cause the data processing hardware (510) to perform operations including: obtaining a pre-trained machine learning model (210O) from a remote system (150); receiving input data (221) captured by the system (100); processing the input data (221) using a machine learning model (210O) on the device corresponding to the pre-trained machine learning model (210O) to generate a plurality of predicted outputs (222); obtaining performance data (250) representing one or more performance characteristics of the machine learning model (210O) on the device, the one or more performance characteristics characterizing the performance of the machine learning model (210O) on the device based on the plurality of predicted outputs (222); Generate one or more performance metrics (302) for the on-device machine learning model (210O) using the performance data (250) without exposing the content of the input data (221) or the multiple predicted outputs (222) to the remote system (150); and Transmit the one or more performance metrics (302) to the remote system (150).
18. The system (100) according to claim 17, wherein, The performance data (250) includes the difference between the multiple predicted outputs (222) and one or more user corrections to the multiple predicted outputs (222).
19. The system (100) according to claim 18, wherein: The difference includes the number of edits to the multiple predicted outputs (222) based on the one or more user corrections; and Generating the one or more performance metrics (302) includes determining an edit rate based on the number of edits.
20. The system (100) according to claim 18 or 19, wherein: For each specific predicted output (222) among the multiple predicted outputs (222), the difference includes an indication of whether the user corrected the specific predicted output (222); and Generating the one or more performance metrics (302) includes determining the incidence of user corrections based on the indication.
21. The system (100) according to any one of claims 17 to 20, wherein, The performance data (250) includes the prediction likelihood determined by the on-device machine learning model (210O) when generating the multiple predicted outputs (222).
22. The system (100) according to any one of claims 17 to 21, wherein, The performance data (250) includes at least one of the following: The amount of time to generate the multiple predicted outputs (222); The memory usage to generate the multiple predicted outputs (222); or The failure of the machine learning system to execute the on-device machine learning model (210O).
23. The system (100) according to any one of claims 17 to 22, wherein, Obtaining the performance data (250) includes obtaining the performance data (250) for multiple time steps.
24. The system (100) according to any one of claims 17 to 23, wherein: The operation further includes updating the on-device machine learning model (210O) over time based on the multiple predicted outputs (222) and one or more user corrections to the multiple predicted outputs (222); and The performance data (250) includes the prediction accuracy of the on-device machine learning model (210O) over time when the system (100) updates the on-device machine learning model (210O).
25. The system (100) according to claim 24, wherein, The performance data (250) further includes the number of parameter values of the machine learning model (210O) on the device over time.
26. The system (100) according to claim 24 or 25, wherein, the prediction accuracy includes an indication that an update to the machine learning model (210O) on the device results in the machine learning model (210O) on the device learning insufficiently or overly in response to user correction.
27. The system (100) according to any one of claims 17 to 26, wherein the operation further includes: when the system updates the machine learning model (210O) on the device, storing a snapshot of the machine learning model (210O) on the device; and restoring the machine learning model (210O) on the device to the stored snapshot based on one or more of the performance metrics (302).
28. The system (100) according to any one of claims 17 to 27, wherein, obtaining the machine learning model (210O) on the device includes obtaining a specific metric definition (304) for each specific performance metric (302) among the one or more performance metrics (302), the specific metric definition including: an indication of the performance data (250) related to the specific performance metric (302) to be obtained; and logic (342) for generating the specific performance metric (302).
29. The system (100) according to claim 28, wherein, the specific metric definition (304) further includes logic (342) for taking an action based on the value of the specific performance metric (302).
30. The system (100) according to claim 29, wherein, the logic (342) for taking the action causes the data processing hardware (510) to perform at least one of the following: transmitting the specific performance metric (302) to the remote system (150); restoring the machine learning model (210O) on the device to a previous state; disabling the machine learning model (210O) on the device; stopping the update of the machine learning model (210O) on the device; or replacing the machine learning model (210O) on the device with a different machine learning model (210O).
31. The system (100) according to any one of claims 28 to 30, wherein, the specific metric definition (304) is generated by a developer who: generated the pre-trained machine learning model (210T); deployed the pre-trained machine learning model (210T) to the system (100) and one or more other user devices (110) via the remote system (150); received the one or more performance metrics (302) from the system (100) and the one or more other user devices (110) via the remote system (150); and Analyze the one or more performance metrics (302) from the system (100) and the one or more other user devices (110) to evaluate the operation of the pre-trained machine learning model (210T).
32. The system (100) according to any one of claims 17 to 31, wherein, the operation further comprises: storing the one or more performance metrics (302) on the system (100); and transmitting the one or more performance metrics (302) to the remote system (150) based on at least one of: a periodic schedule, a received request, a value of a specific performance metric (302) among the one or more performance metrics (302), or an error condition.