Tiny Machine Learning
By generating micro machine learning variants, the problem of high computational overhead of automatic selection of machine learning algorithms in the prior art is solved, and the efficiency and accuracy of algorithm selection are improved, achieving efficient computing resource utilization.
Patent Information
- Application Number
- CN201980081260.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-10-19
- Filing Date
- 2019-10-17
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2039-10-17
AI Technical Summary
When the prior art automatically explores and selects machine learning algorithms, the computing overhead is high, making it difficult to effectively select the best-performing algorithm. The accuracy of the landmark algorithm is low, and it is impossible to accurately indicate the variant performance of the machine learning algorithm.
By generating a micro machine learning variant (micro ML variant) of the machine learning algorithm, this variant is similar enough to the reference variant in terms of prediction results, but has a low computational cost. The micro ML variant reduces the computational cost of training by adjusting the hyperparameter value while maintaining high performance accuracy.
It realizes the calculation cost of machine learning algorithm training while maintaining performance accuracy, saves computing resources, and improves the efficiency of automatic algorithm selection.
Smart Images

Figure CN113168575B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of machine learning and deep learning in artificial intelligence, and particularly to mini-machine learning algorithms. Background Art
[0002] The methods described in this section are methods that can be adopted, but not necessarily methods that have been previously conceived or adopted. Therefore, unless otherwise stated, no method in this section should be considered eligible as prior art solely due to its inclusion in this section.
[0003] Machine learning techniques are now used in a wide variety of applications. Decades of research have created a variety of algorithms and techniques that can be applied to these applications. Selecting the best algorithm for an application can be difficult and resource-intensive. For example, a classification task can be accomplished by several algorithms such as support vector machines (SVMs), random forests, decision trees, artificial neural networks, etc. Each of these algorithms has many variants and configurations and performs differently on different data sets. Selecting the best algorithm is usually a manual task performed by data scientists or machine learning experts with years of experience.
[0004] On the other hand, if various machine learning techniques are automatically explored on data, then there will be a large amount of computational overhead such that it may even delay the entry of machine learning-based products into the market. Therefore, it may not be feasible to automatically train and validate each machine learning algorithm to find the algorithm with the best performance.
[0005] To avoid exploring every machine learning technique, automatic methods of selective training usually end up using a single regressor / classifier to predict algorithm performance without training or validating the algorithm. For example, such a regressor / classifier can be a "landmark" algorithm that is applied to a data set (or a sample thereof) to predict the performance of an unrelated machine learning algorithm. Landmark algorithms such as naive Bayes-type algorithms, k-nearest neighbor algorithms with one neighbor, decision stumps, or algorithms based on principal component analysis (PCA) can be used to predict the performance of neural networks, decision trees, or neural network algorithms.
[0006] However, since the machine learning algorithm using the landmark algorithm is different from the landmark algorithm itself, the prediction accuracy is relatively low. Therefore, even if a machine learning algorithm can be selected based on the landmark, it may not actually represent the best machine learning algorithm for predicting the result.
[0007] In addition, the landmark algorithm approach also does not consider variants of the same algorithm, which can significantly affect the performance and behavior of the algorithm. Whether using automatic algorithm selection or any other method to select a specific algorithm, the performance of the selected algorithm can vary based on variants within the same algorithm. Since the landmark algorithm is different from the machine learning algorithm to be selected, it cannot indicate which variant of a specific machine learning algorithm will result in higher or lower accuracy. The landmark algorithm has no information about a specific machine learning algorithm and is not associated with a specific machine learning algorithm, so it cannot provide any substantial information about variants of a specific machine algorithm. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] In the drawings of certain embodiments, where like reference numerals refer to corresponding parts throughout the figures:
[0009] Figure 1 is a flowchart depicting a process for generating metrics for a reference variant of a machine learning (ML) algorithm in an embodiment;
[0010] Figure 2 is a flowchart depicting a process for generating a micro machine learning (micro-ML) variant of an ML algorithm based on a reference variant of the ML algorithm in an embodiment;
[0011] Figure 3A is a graph depicting the performance scores of a reference variant of an ML algorithm and a micro-ML variant of the ML algorithm for multiple training data sets in an embodiment;
[0012] Figure 3B is a graph depicting the percentage acceleration of a micro-ML variant relative to a reference variant for a training data set in an embodiment;
[0013] Figure 4 is a flowchart depicting a process for generating (one or more) meta-features based on a micro-ML variant for the data set in an embodiment.
[0014] Figure 5 is a block diagram of a basic software system in one or more embodiments;
[0015] Figure 6 is a block diagram of a computer system on which embodiments of the present invention can be implemented. DETAILED DESCRIPTION
[0016] In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the present invention. However, it will be apparent that the present invention may be practiced without these specific details. In other instances, structures and devices are shown in block diagram form to avoid unnecessarily obscuring the present invention.
[0017] Training a Machine Learning Model
[0018] Machine learning techniques involve applying a machine learning algorithm to a training data set for which the (one or more) results are known and which has initialization parameters, the values of which are modified in each training iteration to more accurately produce the (one or more) known results (referred to herein as "(one or more) labels"). Based on such (one or more) applications, these techniques generate a machine learning model with known parameters. Thus, a machine learning model includes a model data representation or a model artifact. The model artifact includes parameter values that are applied by the machine learning algorithm to the input to generate a predicted output. Training a machine learning model requires determining the parameter values of the model artifact. The structure and organization of the parameter values depend on the machine learning algorithm.
[0019] Thus, the term "machine learning algorithm" (or simply "algorithm") as used herein refers to a process or set of rules to be followed in computing, where the model artifact for one or more parameters used in the computation is unknown, while the term "machine learning model" (or simply "model") as used herein refers to a process or set of rules to be followed in computing, where the model artifact for one or more parameters is known and has been derived based on training with a corresponding machine learning algorithm using one or more training data sets. Once trained, an input is applied to the machine learning model for prediction, which may also be referred to herein as a predicted result or output.
[0020] In supervised training, a supervised training algorithm uses training data to train a machine learning model. The training data includes inputs and "known" outputs, labels. In an embodiment, the supervised training algorithm is an iterative process. In each iteration, the machine learning algorithm applies the model artifact and the inputs to generate a predicted output. A cost function is used to compute the error or variance between the predicted output and the known output. In effect, the output of the cost function indicates the accuracy of the machine learning model based on a particular state of the model artifact in the iteration. By applying an optimization algorithm based on the cost function, the parameter values of the model artifact can be adjusted. These iterations can be repeated until a desired accuracy is achieved or some other criterion is satisfied.
[0021] In an embodiment, to iteratively train an algorithm to generate a training model, a training dataset can be arranged such that each row of the dataset is input into a machine learning algorithm, and the corresponding actual result, the label value, of that row is further stored. For example, each row of an adult income dataset represents a specific adult whose result is known, such as whether the total income of that adult exceeds $500,000. Each column of the adult training dataset contains a numerical representation of a specific adult characteristic (e.g., whether the adult has a college degree, the age of the adult...), based on which, when trained, the algorithm can accurately predict whether the total income of any adult (even an adult not yet described by the training dataset) exceeds $500,000.
[0022] The row values of the training dataset can be provided as input to a machine learning algorithm and can be modified based on one or more parameters of the algorithm to produce a prediction result. The prediction result of the row is compared with the label value, and an error value is calculated based on the difference. One or more error values of the batch of rows are used in a statistical aggregation function to calculate the error value of the batch. The term "loss" refers to the error value of a batch of rows.
[0023] At each training iteration, based on the one or more calculated prediction values, the corresponding loss value for that iteration is calculated. For the next training iteration, one or more parameters are modified based on the current loss to reduce the loss. Any number of iterations can be performed on the training dataset to reduce the loss. The training iteration using the training dataset can be stopped when the change in loss between iterations is within a threshold. In other words, when the losses of different iterations are substantially the same, the iteration is stopped.
[0024] After the training iteration, the generated machine learning model includes a machine learning algorithm with a model artifact that has produced the minimum loss.
[0025] For example, a support vector machine (SVM) algorithm can be used to iterate over the above adult income dataset to train an SVM-based model for the adult income dataset. Each row of the adult dataset is provided as input to the SVM algorithm, and the result (prediction result) of the SVM algorithm is compared with the actual result of that row to determine the loss. Based on the loss, the parameters of the SMV are modified. The next row is provided to the SVM algorithm with the modified parameters to produce the prediction result of the next row. This process can be repeated until the difference between the loss values of the previous iteration and the current iteration is below a predefined threshold, or in some embodiments, until the difference between the achieved minimum loss value and the loss of the current iteration is below a predefined threshold.
[0026] Once a machine learning model for a machine learning algorithm is determined, a new dataset with unknown results can be used as input to the model to calculate the prediction result(s) of the new dataset.
[0027] In software implementation, when a machine learning model is said to receive an input, be executed, and / or generate an output or prediction, a computer system executing the machine learning algorithm processes the input by applying the model artifact to generate a predicted output. The computer system processing executes the machine learning algorithm by executing software configured to cause the algorithm to execute.
[0028] Machine learning algorithms and domains
[0029] A machine learning algorithm can be selected based on the domain of the problem and the expected type of result required for the problem. Non-limiting examples of algorithm result types can be discrete values for problems in the classification domain, continuous values for problems in the regression domain, or anomaly detection problems in the clustering domain.
[0030] However, even for a specific domain, there are many algorithms to choose from to select the most accurate algorithm to solve a given problem. As a non-limiting example, in the classification domain, a support vector machine (SVM), random forest (RF), decision tree (DT), Bayesian network (BN), a stochastic algorithm such as a genetic algorithm (GA), or a connectionist topology such as an artificial neural network (ANN) can be used.
[0031] The implementation of machine learning can rely on matrices, symbolic models, and hierarchical and / or associative data structures. Parameterized (i.e., configurable) implementations of the best kinds of machine learning algorithms can be found in open-source libraries such as Google's TensorFlow for Python and C++ or the Georgia Institute of Technology's MLPack for C++. Shogun is an open-source C++ ML library with adapters for several programming languages including C#, Ruby, Lua, Java, MatLab, R, and Python.
[0032] Hyperparameters, cross-validation, and algorithm selection
[0033] One type of machine algorithm can have an infinite number of variants based on one or more hyperparameters. The term "hyperparameter" refers to a parameter in the model artifact that is set before the training of the machine algorithm model and is not modified during the training of the model. In other words, a hyperparameter is a constant value that affects (or controls) the generated trained model independent of the training data set. A machine learning model with a model artifact having only hyperparameter values set is referred to herein as a "variant of a machine learning algorithm" or simply a "variant". Accordingly, during the training of the model, different hyperparameter values of the same type of machine learning algorithm may result in significantly different loss values on the same training data set.
[0034] For example, the SVM machine learning algorithm includes two hyperparameters: "C" and "gamma". The hyperparameter "C" can be set to any value from 10 -3 to 10 5 , while the hyperparameter "gamma" can be set to 10 -5 to 10 3 . Correspondingly, for training the same Adult Income training dataset, there are an infinite number of permutations of the "C" and "gamma" parameters that may produce different loss values.
[0035] Therefore, to select the type of algorithm, or in addition to select the best performing variant of the algorithm, various hyperparameter selection techniques are used to generate a unique set of hyperparameter values. Non-limiting examples of hyperparameter value selection techniques include Bayesian optimization, such as Gaussian processes for hyperparameter value selection, random search, gradient-based search, grid search, manual tuning techniques, techniques based on tree-structured Parzen estimators (TPE).
[0036] Using a unique set of hyperparameter values selected based on one or more of these techniques, each machine learning algorithm variant is trained on the training dataset. The test dataset is used as input to the trained model to calculate the predicted result values. The predicted result values are compared with the corresponding label values to determine the performance score. The performance score can be calculated based on the error rate of the calculated predicted results relative to the corresponding labels. For example, in the classification domain, if only 9,000 out of 10,000 inputs of the model match the label of the input, then the performance score is calculated as 90%. In non-classification domains, the performance score can also be based on the statistical summary of the difference between the label value and the predicted result value.
[0037] The term "experiment" as used herein refers to training a machine learning algorithm using a unique set of hyperparameter values and testing the machine learning algorithm using at least one test dataset. In an embodiment, cross-validation techniques such as k-fold cross-validation are used to create many pairs of training datasets and test datasets from the original training dataset. Each pair of datasets together contains the original training dataset, but the pair of datasets partitions the original dataset in different ways between the training dataset and the test dataset. For each pair of datasets, the training dataset is used to train the model based on the selected set of hyperparameters, and the corresponding test dataset is used to calculate the predicted result values using the trained model. Based on inputting the test dataset into the trained machine learning model, the performance score for that pair (or fold) is calculated. If there is more than one pair (i.e., fold), then the performance scores are statistically summarized (e.g., mean, average, minimum, maximum) to produce the final performance score for that variant of the machine learning algorithm.
[0038] Each trial is computationally expensive because it involves multiple training iterations for variants of a machine algorithm to generate a performance score for a unique set of hyperparameter values of the machine learning algorithm. Thus, reducing the number of trials can significantly reduce the computational resources (e.g., processor time and cycles) required for tuning.
[0039] In addition, since performance scores are generated to select the most accurate algorithm variant, the more precise the performance scores themselves are, the more precise the relative accuracy of the predictions of the generated model will be compared to other variants. In fact, once a machine learning algorithm and its variant based on hyperparameter values are selected, the algorithm variant is applied to the full training dataset using the techniques described above to train a machine model. It is expected that this generated machine learning model will predict results more accurately than machine learning models of any other variant of the algorithm.
[0040] The accuracy of the performance scores themselves depends on how much computational resources are spent on tuning the hyperparameters of the algorithm. Computational resources may be wasted on testing sets of hyperparameter values that cannot produce the desired accuracy of the final model.
[0041] Similarly, tuning those hyperparameters for a type of algorithm that is most likely to be less accurate than another type of algorithm may cost fewer (or no) computational resources. Accordingly, for the hyperparameters of a discounted algorithm, the number of trials can be reduced or eliminated, thus significantly improving the performance of the computer system.
[0042] Overview of Micro Machine Learning Algorithms
[0043] The methods herein describe generating variants of a machine learning algorithm that predict the performance of a computationally more expensive variant of the same type of machine learning algorithm for training. The term "reference variant of the algorithm" or simply "reference variant" as used herein refers to a variant of the algorithm that is computationally more expensive for training than another variant of the same type of machine learning algorithm. The term "micro machine learning algorithm variant" (or "micro ML variant") refers to a computationally less expensive variant of a machine learning algorithm for training.
[0044] In an embodiment, the system evaluates a machine learning algorithm to determine the computational cost and accuracy impact of different unique sets of hyperparameter values. The unique sets of hyperparameter values that are determined to produce an acceptable accuracy while reducing the computational cost of training the algorithm are included in the micro ML variant.
[0045] In an embodiment, the accuracy and computational cost are measured compared to a reference variant of the algorithm. The reference variant is selected as a baseline variant of a machine learning algorithm, which is modified to generate a TinyML variant. The TinyML variant is generated by iteratively evaluating variants of the reference variant, and in each iteration, at least one hyperparameter value of the reference variant is modified. The modified algorithm is evaluated using a training dataset to determine a performance score and to determine the computational cost for training the modified algorithm using the training dataset. Based on the computational cost and the performance score, the system determines whether to select the modified hyperparameter value for the TinyML variant of the algorithm - the smaller the computational cost, the more likely the hyperparameter value is to be selected.
[0046] In an embodiment, the hyperparameter value with the minimum computational cost is selected. The selection may also be based on the performance score meeting a performance criterion for the TinyML variant. The performance criterion may be based on the performance score of the reference variant.
[0047] Generate a reference variant of a machine learning algorithm
[0048] Figure 1 is a flowchart depicting the process of generating metrics for a reference variant in an embodiment. In an embodiment, at step 100, a computer system selects a reference variant with a predetermined set of hyperparameters for a specific domain. For example, based on previous experiments, it has been shown that for classification domain problems involving three or fewer input types, specific "C" and "gamma" values yield the best performance. This variant of the SVM machine learning algorithm is then selected as the reference variant to generate a TinyML variant. At step 110, a training dataset is selected for cross - validation of the reference variant, and at step 120, cross - validation of the reference variant is performed on the selected dataset. In some embodiments, the reference variant may already be associated with pre - recorded performance and cost metrics for the dataset, which are used to generate the TinyML variant. For example, the TinyML variant will be generated based on one or more datasets from OpenML.org, and the reference variant has been cross - validated and trained using the same datasets.
[0049] In an embodiment, to obtain the performance score of the reference variant, the reference variant is cross - validated using the selected training dataset, which is also used to generate the TinyML variant of the algorithm. The performance score can be calculated by statistically aggregating the individual performance scores calculated for each dataset fold validation. At step 125, if there are other training datasets, then cross - validation is similarly performed using each other training dataset to generate a corresponding performance score for each other training dataset.
[0050] In another embodiment, reference variants are generated by tuning a particular type of algorithm on one or more training datasets in a particular domain. Using the techniques described above, different unique sets of hyperparameters can be used to perform cross-validation on one or more datasets in a particular domain. The trial that produces the highest aggregate performance score for the (one or more) training datasets is selected as the reference variant.
[0051] For example, continuing Figure 1 , at step 100, one variant of a machine learning algorithm is selected, but this algorithm has not yet been designated as the reference variant. At steps 110 - 125, cross-validation is performed on the training datasets using the selected variant of the algorithm. At step 130, the system uses hyperparameter selection techniques and the performance scores from the cross-validation to determine whether to perform further tuning of the hyperparameter values of the reference variant. If so, then at step 133, the system generates a new unique set of hyperparameter values according to the techniques described above, thereby generating a new variant of the algorithm. Steps 100 - 125 are repeated for the new variant until at step 130, the system determines that no further tuning is needed based on the generated performance scores. At step 135, the machine learning algorithm with the best performance score is selected as the reference variant.
[0052] Additionally or alternatively, at step 140, regardless of the technique for selecting the reference variant, a computational cost metric is generated for the reference variant. In one embodiment, this computational cost metric can be generated by measuring the duration used to generate the machine learning model by training the reference variant on the training datasets. The duration can be measured based on training using predefined computational resources.
[0053] Other embodiments of generating the computational cost metric, as a supplement or alternative to the duration measurement, include measurements based on one or more of memory consumption, processor utilization, or I / O activity measurements.
[0054] The computational cost metric for each training dataset can be recorded separately. For example, continuing Figure 1 , at step 145, a tuple is generated for each training dataset that contains the computational cost metric and the performance score for that dataset.
[0055] Generating a miniature machine learning variant of the algorithm
[0056] In an embodiment, a miniature ML variant is generated based on the reference variant. The (one or more) hyperparameter values of the reference variant are modified to produce a miniature ML variant that is similar enough to the reference variant in terms of prediction results but has a lower computational cost at the same time. In other words, the miniature ML variant is a variant of the algorithm that follows the performance of the reference variant but is computationally cheaper when trained on the (one or more) datasets.
[0057] Figure 2 It is a flowchart depicting the process of generating a micro ML variant based on a reference variant in an embodiment. At step 200, an accuracy criterion is selected. The term "accuracy criterion" refers to a criterion for determining an acceptable measure of the deviation between the accuracy of the micro ML variant and the accuracy of the reference variant. In an embodiment, the accuracy criterion is based on the difference between the performance scores of the reference variant and the micro ML variant, and the difference between the (one or more) performance scores of the reference variant and the corresponding performance scores of the micro ML variant. For example, the accuracy criterion may indicate that the difference in performance scores must be within a specific percentage of each other, or below a specific absolute value.
[0058] At step 205, hyperparameters are selected from the reference variant to be modified to generate a new variant of the algorithm. At step 210, the values of the selected hyperparameters are modified in the reference variant, thereby generating a new variant of the algorithm.
[0059] In an embodiment, compared to using the reference variant, the hyperparameter values are modified so that less computational resources are used for training. For example, reducing the values of hyperparameters for multiple layers in a neural network algorithm makes the training of the neural network computationally less expensive. Similarly, reducing the hyperparameter value representing the number of units in the first layer of a neural network also achieves a reduction in the computational cost of the neural network algorithm.
[0060] In an embodiment where the hyperparameter values are independent of the computational cost of the training model, any of the above hyperparameter value selection techniques (e.g., Bayesian, random search, gradient descent) are used to select new hyperparameter values.
[0061] At step 215, a training dataset is selected from a plurality of training datasets to perform cross - validation on the new variant of the algorithm including the new hyperparameter values. The plurality of training sets available for selection at step 215 are those for which performance scores exist or for which performance scores will be calculated for the reference variant using the above - mentioned techniques. At step 220, the system performs cross - validation on the selected training dataset, thereby calculating the performance score of the selected training dataset.
[0062] At step 225, the calculated performance score for the selected training dataset is compared with the performance score of the reference variant for the same training dataset. If the performance score meets the accuracy criterion, then the process proceeds to step 230. If the performance score does not meet the accuracy criterion, then the process returns to step 205 or step 210. In other words, if the performance score of the selected variant of the algorithm is similar to the performance score of the reference variant for the same dataset, then the new variant is eligible as a candidate for the micro ML variant of the algorithm.
[0063] For example, an accuracy criterion may specify a maximum 10% deviation in measurements for a modified algorithm. In such an example, the difference in the performance score of the computational performance score and the reference variant is calculated and compared to the 10% accuracy criterion. If the calculated ratio does not exceed 10%, then the accuracy criterion is met. Otherwise, the accuracy criterion has failed for the variant algorithm.
[0064] In embodiments where the accuracy criterion has failed, the process determines whether to select a new value for the hyperparameter at step 210 or a new hyperparameter itself at step 205. This determination is based on whether the value selection of the hyperparameter value depends on the performance score. It is expected that hyperparameter value selection techniques based on predicted computational costs will in turn affect the performance score.
[0065] For example, if for a neural network-based reference variant, the number of layers in the network hyperparameters decreases in each iteration, then once the accuracy criterion is not met, the process will not perform experiments to set even lower values for that hyperparameter. The lower value of the hyperparameter is expected to produce a more inaccurate model and thus ensure failure to pass the accuracy criterion. Accordingly, the process can transition to another hyperparameter at step 205.
[0066] In other words, if the current hyperparameter value is selected at step 210 in such a way that it may result in a lower performance score than the previous hyperparameter value for the same hyperparameter, then similarly, the next hyperparameter value will result in an even lower performance score than the current hyperparameter value. Accordingly, for the next selection of the hyperparameter value, failure to pass the accuracy criterion is also expected. In such a case, to save computational resources, the process transitions to step 205 to select a different hyperparameter for the modified reference variant.
[0067] For a modified reference variant that meets the accuracy criterion at step 225, the process transitions to step 230. At step 230, the modified reference variant is trained on the selected training dataset, and a cost metric for that training is followed. For example, the duration required to train the corresponding model on a defined computational resource configuration can be recorded. In one embodiment, the process can determine the type of cost metric that exists for the reference variant and measure the same cost metric for the modified algorithm.
[0068] At step 235, the (one or more) cost metrics of the reference variant and the modified reference variant are compared for the selected training dataset. If the cost metric of the modified reference variant indicates that the computational cost of training the model is lower than that of the reference variant, then the selected hyperparameter value continues to qualify for the TinyML variant. For example, if the previously described training duration metric of the modified reference variant for the selected training dataset is less than the same metric of the reference variant, then the selected hyperparameter value continues to qualify for the TinyML variant of the algorithm.
[0069] In an embodiment, the modified reference variant is selected only if the (one or more) cost metrics for each training dataset indicate less resource consumption. In such an embodiment, the (one or more) cost metrics of the reference variant for each training dataset are compared with the (one or more) cost metrics of the modified reference variant for the corresponding dataset. At step 235, if the cost metric of the modified reference variant indicates that even for a single dataset, more computational resources are consumed for training, then the hyperparameter value is rejected for the TinyML variant.
[0070] In a related embodiment, the modified reference variant must substantially improve the computational cost for training the dataset to be selected. In such an embodiment, if the cost metric fails to indicate an improvement in the threshold amount of computational resources for training the dataset, then the modified reference variant is rejected at step 235. For example, the threshold can be defined as an improvement of 50% or more over the cost of the reference variant. The cost metric of the reference variant is divided by the cost metric of the modified algorithm, and if the division is equal to or greater than 50% by percentage, then the hyperparameter value qualifies for the TinyML variant.
[0071] In another embodiment, the modified reference variant may not be rejected at step 235 unless the cost metric for each and every selected training dataset indicates a lower computational cost for the modified reference variant. In such an embodiment, at step 240, if for a threshold number of datasets, the algorithm based on the modified hyperparameter value is less costly than the reference variant, then the modified hyperparameter value can qualify. The threshold number can be determined based on a percentage of the total number of training datasets available for the process. For example, if the threshold percentage is 90%, then if for 9 out of 10 training datasets, the algorithm based on the modified hyperparameter value is less costly, then the hyperparameter value qualifies for the TinyML variant.
[0072] At step 245, if more than one hyperparameter value qualifies for the hyperparameter, then the hyperparameter value for which the modified algorithm has already produced the lowest cost metric (lowest computational resource cost) is selected for the microML variant. To determine the hyperparameter value to be selected for a hyperparameter, the cost metrics for each hyperparameter are statistically aggregated, and the hyperparameter value with the best performance statistically aggregated cost metric is selected (e.g., the lowest aggregated cost metric value).
[0073] In an embodiment, if no modified value for a hyperparameter qualifies for the microML variant, then the original hyperparameter value of the reference variant is used for the specific hyperparameter in the microML algorithm.
[0074] Figure 3A is a graph depicting the performance scores of the reference variant and the microML variant of the algorithm for multiple training datasets in an embodiment. Each point on the graph corresponds to the performance score of a specific dataset in the OpenML.org dataset. The line for the reference dataset overlaps with the line for the microML variant, indicating that the microML variant closely follows the performance of the reference variant.
[0075] Figure 3B is depicting in an embodiment Figure 3A of the same microML variant relative to Figure 3A of the same reference variant for the percentage of acceleration for the training training dataset. Each point is the percentage of the difference in the duration of training a specific dataset using the reference variant and the corresponding microML variant relative to the duration of the reference variant for training the same dataset. As shown, the microML variant represents a significant acceleration in training the model relative to the reference model for the OpenML.org dataset without losing the accuracy of the reference variant, as depicted by the performance scores in Figure 3A .
[0076] Use of the micro machine learning algorithm
[0077] In an embodiment, since the micro machine learning variant of the algorithm has a lower training cost than the reference variant while following the performance of the reference variant, the micro machine learning of the algorithm is used to determine the performance of the reference variant on a dataset. For example, instead of cross - validating multiple different reference variants for a new training dataset to determine which reference variant will perform best for that new dataset, the corresponding microML variant is used for this determination. The microML variant is cross - validated instead of the corresponding reference variant, which saves a large amount of computational resources. Based on the calculated performance scores in the cross - validation from the microML, the reference variant with the highest score of its microML variant is selected. Accordingly, the selection process of the reference variant is performed using the corresponding microML variant to save computational resources.
[0078] In another embodiment, the performance scores of the microML variants on the training datasets are used to train a meta-model for hyperparameter value selection or machine algorithm selection. One or more reference variants are selected, and hyperparameters are tuned for these one or more reference variants or these one or more reference variants are considered for algorithm selection. The generated microML variants corresponding to the reference variants are identified to generate meta-feature values for the training datasets. In an embodiment, the system has recorded the performance scores of the individual training datasets when generating the microML variants, and thus, this performance score can be easily used as a meta-feature when training the meta-model. For training datasets for which the system does not have performance scores for the microML variants, cross-validation of the microML variants will be performed on this training dataset.
[0079] Figure 4 is a flowchart depicting the generation of (one or more) meta-features based on microML variants for a dataset in an embodiment. At step 410, a training dataset is selected for which (one or more) meta-features based on microML variants are to be generated. At step 420, cross-validation of the microML variants is performed on the selected training dataset to generate the performance scores of the microML variants for this training dataset. If, at step 430, there is another microML variant, then cross-validation of this another microML variant is performed on the selected dataset to produce another performance score until cross-validation of all microML variants is performed for the selected training dataset. At step 440, the generated performance scores are designated as the meta-features of the selected dataset.
[0080] In one or more embodiments, the generated meta-features based on microML variants are used to tune hyperparameter values and / or select machine learning algorithms in training the (one or more) meta-models as a supplement or alternative to the existing meta-features of the training datasets.
[0081] One such example is the scalable and efficient distributed auto-tuning of machine learning models and deep learning models. In such an example, a gradient-based auto-tuning algorithm can be used, which includes the following three main steps. First, given a list of parameters and their ranges (which form the parameter space to be searched), start by running (training and evaluating) all combinations of the boundary values of the parameter space in parallel. Second, for each parameter in the parameter space, the (usually wide) input parameter range is reduced in parallel by using a gradient descent method (described later). Finally, using the refined parameter space of all parameters, a simple gradient descent is performed from the best point in the corresponding parameter space for each parameter.
[0082] Using meta - learning for gradient - based automatic hyperparameter optimization for machine learning models and deep - learning models is another example of training one or more meta - models that are used to leverage the generated meta - features based on ML variants to tune hyperparameter values and / or select machine - learning algorithms. In such an example, techniques are provided for optimal initialization of machine - learning algorithm hyperparameters and other predicted value ranges based on dataset meta - features. In an embodiment for each specific hyperparameter of a machine - learning algorithm, the computer invokes a unique trained meta - model for the specific hyperparameter based on an inferred dataset to detect an improved sub - range of possible values for the specific hyperparameter. The machine - learning algorithm is configured based on the improved sub - range of possible values for the hyperparameter. The machine - learning algorithm is invoked to obtain results. In an embodiment, gradient - based search - space reduction (GSSR) finds the optimal value within the improved sub - range of values for a specific hyperparameter. In an embodiment, the meta - model is trained based on dataset meta - features and performance metrics from exploratory sampling of a configured hyperspace (such as via GSSR). In various embodiments, other values are optimized or intelligently predicted based on additional trained / trainable meta - models.
[0083] Using an algorithm - specific neural - network architecture for automatic machine - learning model selection is another example of training one or more meta - models that are used to leverage the generated meta - features based on micro - ML variants to tune hyperparameter values and / or select machine - learning algorithms. In such an example, techniques are provided for optimal selection of machine - learning algorithms based on performance prediction of a trained algorithm - specific regressor. In an embodiment, the computer derives respective meta - feature values from an inferred dataset by deriving a corresponding meta - feature value from the inferred dataset for each meta - feature. For each trainable algorithm and each regression meta - model respectively associated with the algorithm, a corresponding score is calculated by invoking the meta - model based on at least one of the following: a) a respective subset of the meta - feature values, and / or b) hyperparameter values of a respective subset of the hyperparameters of the algorithm. One or more algorithms are selected based on the respective scores. Based on the inferred dataset, one or more algorithms can be invoked to obtain results. In an embodiment, the trained regressor is an artificially - neural network configured differently. In an embodiment, the trained regressor is contained within an algorithm - specific ensemble. Techniques are also provided for optimal training of the regressor and / or the ensemble.
[0084] Using gradient-based auto-tuning for machine learning models and deep learning models is another example of training a meta-model that uses meta-features generated from microML variants to tune hyperparameter values and / or select machine learning algorithms. In such an example, horizontally scalable techniques are provided for efficient configuration of machine learning algorithms to achieve optimal accuracy without the need for a known input. In an embodiment, for each specific hyperparameter that is not a categorical hyperparameter, and for each epoch in an epoch sequence, the computer processes the specific hyperparameter as follows. An epoch explores one hyperparameter based on a computer-generated hyperparameter tuple. For each tuple, a score is calculated based on the tuple. A hyperparameter tuple contains a unique combination of values, each value being included in the current value range of a unique hyperparameter. All values of the hyperparameter tuple belonging to a specific hyperparameter are unique. All values of the hyperparameter tuple belonging to any other hyperparameter remain constant during the epoch, such as the best value of that other hyperparameter so far. The computer narrows the current value range of the specific hyperparameter based on the intersection of a first line based on the scores and a second line based on the scores. Based on the repeatedly narrowed value range of the hyperparameter, the machine learning algorithm is optimally configured. The configured algorithm is invoked to obtain a result, such as pattern recognition or classification between multiple possible patterns.
[0085] In an embodiment, for any new dataset, hyperparameter value selection or machine algorithm model selection is performed by applying the meta-features of the new dataset as input to the trained meta-model. The system performs cross-validation of the microML variant on the new dataset and uses the resulting performance score as the meta-feature value. The meta-feature value of the new dataset, including the meta-feature value of the microML variant(s), is provided as input to the meta-model. The output of the meta-model indicates the best performing machine learning algorithm to be predicted and / or one or more hyperparameter values of the best performing machine learning algorithm to be predicted.
[0086] The microML variant as a meta-feature significantly improves the accuracy of the meta-model since the microML variant follows the performance of the corresponding reference variant. Thus, the meta-learning algorithm has an accurate input that describes the (one or more) dataset performance of various machine learning algorithms and thus produces a more accurate meta-model. Since the cost of training or cross-validating the microML variant on a dataset is lower than that of the reference variant, an improved accuracy of the meta-model can be achieved with only a small computational cost.
[0087] Software Overview
[0088] Figure 5 can be used to control Figure 6Block diagram of the basic software system 500 for the operation of the computing system 600. The software system 500 and its components, including their connections, relationships, and functions, are merely exemplary and are not meant to limit the implementation of the example embodiment(s). Other software systems suitable for implementing the example embodiment(s) may have different components, including components with different connections, relationships, and functions.
[0089] A software system 500 is provided for guiding the operation of the computing system 600. The software system 500, which can be stored in the system memory (RAM) 606 and the storage device (e.g., hard disk or flash memory) 610, includes a kernel or operating system (OS) 510.
[0090] The OS 510 manages the low-level aspects of computer operation, including managing the execution of processes, memory allocation, file input and output (I / O), and device I / O. One or more applications, represented as 502A, 502B, 502C... 502N, can be "loaded" (e.g., transferred from the storage device 610 to the memory 606) for execution by the system 500. Applications or other software intended to be used on the computer system 600 can also be stored as a set of downloadable computer-executable instructions, e.g., for downloading and installing from an Internet location (e.g., a web server, an app store, or another online service).
[0091] The software system 500 includes a graphical user interface (GUI) 515 for receiving user commands and data in a graphical (e.g., "point and click" or "touch gesture") manner. These inputs can in turn be acted upon by the system 500 according to instructions from the operating system 510 and / or the application(s) 502. The GUI 515 is also used to display the operation results from the OS 510 and the application(s) 502, after which the user can provide additional input or terminate the session (e.g., log off).
[0092] The OS 510 can execute directly on the bare hardware 520 of the computer system 600 (e.g., the processor(s) 604). Alternatively, a hypervisor or virtual machine monitor (VMM) 530 can be inserted between the bare hardware 520 and the OS 510. In this configuration, the VMM 530 acts as a software "buffer" or virtualization layer between the OS 510 and the bare hardware 520 of the computer system 600.
[0093] The VMM 530 instantiates and runs one or more virtual machine instances ("guest computers"). Each guest machine includes a "guest" operating system (such as OS 510) and one or more applications (such as (one or more) applications 502), which are designed to execute on the guest operating system. The VMM 530 presents a virtual operating platform to the guest operating system and manages the execution of the guest operating system.
[0094] In some cases, the VMM 530 may allow the guest operating system to run as if it were running directly on the bare hardware 520 of the computer system 600. In these cases, the same version of the guest operating system that is configured to execute directly on the bare hardware 520 can also execute on the VMM 530 without modification or reconfiguration. In other words, in some cases, the VMM 530 can provide full hardware and CPU virtualization to the guest operating system.
[0095] In other cases, the guest operating system may be specially designed or configured to execute on the VMM 530 for increased efficiency. In these cases, the guest operating system "knows" that it is executing on the virtual machine monitor. In other words, in some cases, the VMM 530 can provide para-virtualization to the guest operating system.
[0096] A computer system process includes the allocation of hardware processor time and the allocation of memory (physical and / or virtual), the memory being allocated for storing instructions executed by the hardware processor, for storing data generated by the hardware processor executing the instructions, and / or for storing the hardware processor state (e.g., the contents of registers) between allocations of hardware processor time when the computer system process is not running. A computer system process runs under the control of an operating system and can also run under the control of other programs executing on the computer system.
[0097] Multiple threads can run within a process. Each thread also includes an allocation of hardware processing time but shares access to the memory allocated to the process. This memory is used to store the processor contents between allocations when the thread is not running. The term "thread" can also be used to refer to a computer system process when multiple threads are not running.
[0098] Hardware Overview
[0099] According to one embodiment, the techniques described herein are implemented by one or more special-purpose computing devices. The special-purpose computing device can be hard-wired to perform the techniques, or can include digital electronic devices permanently programmed to perform the techniques, such as one or more application-specific integrated circuits (ASICs) or field-programmable gate arrays (FPGAs), or can include one or more general hardware processors programmed to perform the techniques according to program instructions in firmware, memory, other storage devices, or a combination. Such special-purpose computing devices can also combine custom hard-wired logic, ASICs, or FPGAs with custom programming to implement the techniques. The special-purpose computing device can be a desktop computer system, a portable computer system, a handheld device, a networking device, or any other device that combines hard-wired and / or program logic to implement the techniques.
[0100] For example, Figure 6 is a block diagram of a computer system 600 on which embodiments of the invention can be implemented. The computer system 600 includes a bus 602 or other communication mechanism for conveying information, and a hardware processor 604 coupled to the bus 602 for processing information. The hardware processor 604 can be, for example, a general-purpose microprocessor.
[0101] The computer system 600 also includes a main memory 606 coupled to the bus 602 for storing instructions and information to be executed by the processor 604, such as random access memory (RAM) or other dynamic storage device. The main memory 606 can also be used to store temporary variables or other intermediate information during execution of instructions to be executed by the processor 604. When such instructions are stored in a non-transitory storage medium accessible to the processor 604, such instructions cause the computer system 600 to become a special-purpose machine customized to perform the operations specified in the instructions.
[0102] The computer system 600 also includes a read-only memory (ROM) 608 or other static storage device coupled to the bus 602 for storing static information and instructions for the processor 604. A storage device 610, such as a magnetic disk or optical disk, is provided and coupled to the bus 602 for storing information and instructions.
[0103] The computer system 600 can be coupled via a bus 602 to a display 612 for displaying information to a computer user, such as a cathode ray tube (CRT). An input device 614 including alphanumeric keys and other keys is coupled to the bus 602 for transmitting information and command selections to the processor 604. Another type of user input device is a cursor control 616, such as a mouse, trackball, or cursor direction keys, for transmitting direction information and command selections to the processor 604 and for controlling cursor movement on the display 612. Such input devices typically have two degrees of freedom on two axes (a first axis (e.g., x) and a second axis (e.g., y)) to allow the device to specify a position in a plane.
[0104] The computer system 600 can implement the techniques described herein using custom hardwired logic, one or more ASICs or FPGAs, firmware, and / or program logic that in combination with the computer system causes the computer system 600 to be a special-purpose machine or programs the computer system 600 to be a special-purpose machine. According to one embodiment, the techniques herein are performed by the computer system 600 in response to one or more sequences of one or more instructions contained in the main memory 606 being executed by the processor 604. These instructions can be read into the main memory 606 from another storage medium, such as the storage device 610. Execution of the instruction sequences contained in the main memory 606 causes the processor 604 to perform the processing steps described herein. In an alternative embodiment, hardwired circuitry may be used in place of or in combination with software instructions.
[0105] As used herein, the term "storage medium" refers to any non-transitory medium that stores data and / or instructions that cause a machine to operate in a particular manner. Such storage media may include non-volatile media and / or volatile media. Non-volatile media includes, for example, optical or magnetic disks, such as the storage device 610. Volatile media includes dynamic memory, such as the main memory 606. Common forms of storage media include, for example, floppy disks, flexible disks, hard disks, solid state drives, magnetic tape, or any other magnetic data storage medium, CD-ROM, any other optical data storage medium, any physical medium with hole patterns, RAM, PROM, and EPROM, FLASH-EPROM, NVRAM, any other memory chip or cartridge.
[0106] Storage media is distinct from but can be used in conjunction with transmission media. Transmission media participates in transferring information between storage media. For example, transmission media includes coaxial cables, copper wire, and fiber optics, including the wires that comprise the bus 602. Transmission media can also take the form of acoustic or light waves, such as those generated during radio wave and infrared data communications.
[0107] Various forms of media can be involved in carrying one or more sequences of one or more instructions to the processor 604 for execution. For example, the instructions can initially be carried on a magnetic disk or solid state drive of a remote computer. The remote computer can load the instructions into its dynamic memory and send the instructions over a telephone line using a modem. A modem local to the computer system 600 can receive the data on the telephone line and convert the data to an infrared signal using an infrared transmitter. An infrared detector can receive the data carried in the infrared signal, and appropriate circuitry can place the data on the bus 602. The bus 602 carries the data to the main memory 606, and the processor 604 retrieves and executes the instructions from the main memory 606. The instructions received by the main memory 606 can optionally be stored on the storage device 610 before or after being executed by the processor 604.
[0108] The computer system 600 also includes a communication interface 618 coupled to the bus 602. The communication interface 618 provides two-way data communication coupled to a network link 620, where the network link 620 is connected to a local network 622. For example, the communication interface 618 can be an Integrated Services Digital Network (ISDN) card, a cable modem, a satellite modem, or a modem that provides a data communication connection to a corresponding type of telephone line. As another example, the communication interface 618 can be a LAN card that provides a data communication connection to a compatible local area network (LAN). A wireless link can also be implemented. In any such implementation, the communication interface 618 sends and receives electrical, electromagnetic, or optical signals that carry digital data streams representing various types of information.
[0109] The network link 620 typically provides data communication to other data devices through one or more networks. For example, the network link 620 can provide a connection to a main computer 624 or to a data device operated by an Internet service provider (ISP) 626 through the local network 622. The ISP 626 in turn provides data communication services through the global packet data communication network now commonly referred to as the “Internet” 628. Both the local network 622 and the Internet 628 use electrical, electromagnetic, or optical signals that carry digital data streams. Signals through the various networks and signals on the network link 620 and through the communication interface 618 are example forms of transmission media that carry digital data to or from the computer system 600.
[0110] The computer system 600 can send messages and receive data, including program code, via one or more networks, network links 620, and communication interface 618. In an Internet example, the server 630 can transmit the requested code for an application program via the Internet 628, an ISP 626, a local network 622, and communication interface 618.
[0111] The received code can be executed by the processor 604 when it is received and / or stored in the storage device 610 or other non-volatile storage means for later execution.
Claims
1. A computer-implemented method, comprising: selecting a reference variant of a machine learning algorithm, the reference variant of the machine learning algorithm indicating one or more hyperparameter values; modifying at least one hyperparameter to a new hyperparameter value among the one or more hyperparameter values, thereby generating a new variant of the machine learning algorithm; using a training dataset to determine a performance score of the new variant of the machine learning algorithm, the performance score representing the accuracy of a new machine learning model generated by training the new variant of the machine learning algorithm using the training dataset; comparing the performance score of the new variant of the machine learning algorithm for the training dataset with the performance score of the reference variant of the machine learning algorithm for the training dataset to determine at least one deviation among a plurality of deviations of the performance score between the performance score of the new variant of the machine learning algorithm and the performance score of the reference variant of the machine learning algorithm; determining a cost metric of the new variant of the machine learning algorithm by measuring the use of one or more of memory consumption, processor utilization, or I / O activity measurements when training the new variant of the machine learning algorithm on the training dataset; determining whether to use the new variant of the machine learning algorithm as a computationally less costly micro-ML variant for determining the performance of the reference variant for a new dataset based on the cost metric of the new variant of the machine learning algorithm and the plurality of deviations of the performance score that at least partially satisfy one or more accuracy criteria.
2. The method according to claim 1, further comprising: determining that the new variant of the machine learning algorithm satisfies the one or more accuracy criteria based on the performance score of the reference variant of the machine learning algorithm by comparing the performance score of the new variant of the machine learning algorithm for the training dataset with the performance score of the reference variant of the machine learning algorithm for the training dataset; determining to select the new hyperparameter value for the micro-ML variant of the machine learning algorithm based on the cost metric of the new variant of the machine learning algorithm.
3. The method according to claim 1, wherein determining the performance score of the new variant of the machine learning algorithm using the training dataset further comprises performing cross-validation of the new variant of the machine learning algorithm on the training dataset.
4. The method according to claim 1, further comprising determining the performance score of the reference variant by performing cross-validation of the reference variant on the training dataset.
5. The method according to claim 1, further comprising generating a reference variant by: selecting a unique set of hyperparameter values from a plurality of unique sets of hyperparameter values for the machine learning algorithm; performing cross-validation of the machine learning algorithm on one or more training datasets; determining whether to select the unique set of hyperparameter values for the reference variant based on performing cross-validation of the machine learning algorithm.
6. The method according to claim 5, further comprising selecting the unique set of hyperparameter values from the plurality of unique sets of hyperparameter values based on one of: Bayesian optimization, random search, gradient-based search, grid search, or selection based on the tree-structured Parzen estimator TPE.
7. The method according to claim 1, further comprising: determining a cost metric of a reference variant by measuring the usage of computing resources for training the reference variant on the training dataset; comparing the cost metric of a new variant of the machine learning algorithm with the cost metric of the reference variant; determining that the cost metric of the new variant of the machine learning algorithm is lower than the cost metric of the reference variant based on comparing the cost metric of the new variant of the machine learning algorithm with the cost metric of the reference variant; qualifying the new hyperparameter value for a TinyML variant of the machine learning algorithm based on determining that the cost metric of the new variant of the machine learning algorithm is lower than the cost metric of the reference variant.
8. The method according to claim 1, further comprising: qualifying the new hyperparameter value for a TinyML variant of the machine learning algorithm based on the cost metric of the new variant of the machine learning algorithm and by comparing the performance score of the new variant of the machine learning algorithm on the training dataset with the performance score of the reference variant of the machine learning algorithm on the training dataset; modifying the at least one hyperparameter from an original hyperparameter value to another hyperparameter value to generate another machine learning algorithm with the at least one hyperparameter having the another hyperparameter value; comparing the performance score of the another machine learning algorithm on the training dataset with the performance score of the reference variant of the machine learning algorithm on the training dataset; qualifying the another hyperparameter value for a TinyML variant of the machine learning algorithm based on the cost metric of the another algorithm and by comparing the performance score of the another machine learning algorithm with the performance score of the reference variant of the machine learning algorithm; determining whether to select the new hyperparameter value or the another hyperparameter value for a TinyML variant of the machine learning algorithm based on the cost metric of the another algorithm and the cost metric of the new variant of the machine learning algorithm.
9. The method according to claim 1, wherein the new hyperparameter value is based on a previous hyperparameter value of a previous machine learning algorithm generated from a reference variant of the machine learning algorithm.
10. The method according to claim 1, further comprising: determining that the new variant of the machine learning algorithm does not meet the one or more accuracy criteria based on comparing the performance score of the new variant of the machine learning algorithm on the training dataset with the performance score of the reference variant of the machine learning algorithm on the training dataset and based on the performance score of the reference variant of the machine learning algorithm; determining not to select the new hyperparameter value for a TinyML variant of the machine learning algorithm based on determining that the new variant of the machine learning algorithm does not meet the one or more accuracy criteria; modifying the at least one hyperparameter from the original hyperparameter value to a next hyperparameter value to generate a next machine learning algorithm from the reference variant of the machine learning algorithm with the at least one hyperparameter having the next hyperparameter value; determining that the next machine learning algorithm is less accurate than the new variant of the machine learning algorithm without performing cross-validation of the next machine learning algorithm on the training dataset.
11. The method according to claim 1, wherein the reference variant of the machine learning algorithm is one of a plurality of reference variants of the machine learning algorithm, and corresponding plural modified, less costly variants of the reference variant of the machine learning algorithm are generated for the plurality of reference variants of the machine learning algorithm, the corresponding plural modified variants including new variants of the machine learning algorithm, and the method further comprises: receiving a request for determining the expected performance of the plurality of reference variants against a specific training dataset; for each reference variant of the plurality of reference variants of the machine learning algorithm, selecting a corresponding modified variant; performing cross-validation of the corresponding modified variant, thereby generating a corresponding specific performance score for each reference variant; based on the corresponding specific performance scores, determining whether each reference variant of the plurality of reference variants yields the highest accuracy when trained by the specific training dataset.
12. One or more non-transitory storage media storing instructions which, when executed by one or more hardware processors, cause the method according to any one of claims 1-11 to be performed.
13. An apparatus comprising components for performing the method according to any one of claims 1-11.
14. A computer system, comprising: a hardware processor; and a memory coupled to the hardware processor and including instructions stored thereon which, when executed by the hardware processor, cause the hardware processor to perform the method according to any one of claims 1-11.
15. A computer program product comprising instructions which, when executed by a computer, cause the computer to perform the method according to any one of claims 1-11.