Hyper-parameter adjustment

By using historical experimental statistics and performance attribution statistics to allocate weights to hyperparameters, combined with Bayesian update method, we quickly find a set of hyperparameters that meet the adjustment conditions, solving the problem of low efficiency of hyperparameter adjustment in the existing technology, and achieving efficient hyperparameter optimization.

CN120344980APending Publication Date: 2025-07-18MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480005501.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-02-14
Filing Date
2024-02-08
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The prior art has high computational overhead and requires manual management when adjusting the hyperparameters of machine learning models, making it difficult to efficiently find the optimal set of hyperparameters within a specified computing budget.

Method used

By assigning weights to each hyperparameter using historical experimental statistics and performance attribution statistics, the hyperparameters are updated in a series of experiments using Bayesian update method, and the set of hyperparameters that meet the adjustment conditions are selected.

Benefits of technology

While coordinating the calculation budget, it reduces computational overhead and manual intervention, quickly finds a set of hyperparameters that meet performance goals, and improves the efficiency of hyperparameter adjustment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120344980A_ABST
    Figure CN120344980A_ABST
Patent Text Reader

Abstract

The hyper-parameter adjustment system generates, for each hyper-parameter, performance attribution statistics corresponding to the evaluation metrics of the machine learning model based on historical experimental statistics for the evaluation metrics and the machine learning model. The hyper-parameter adjustment system assigns a weight to each hyper-parameter based on performance attribution statistics of the hyper-parameter. In a series of experiments, the hyper-parameter adjustment system updates the hyper-parameters based on a weight assigned at each hyper-parameter, and selects a set of hyper-parameters for the machine learning model from one of the experiments, where the set of hyper-parameters produces recorded values of evaluation metrics that meet adjustment conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Generally, hyperparameters (e.g., those related to architecture complexity and algorithm hyperparameters) are adjustable parameters that affect the performance of a machine learning model. Compared with the internal parameters of the model being trained during the training process, such as the coefficients (or weights) of linear and logistic regression models, the weights and biases of neural networks, and the cluster centers in clustering, hyperparameters define the structure and algorithmic characteristics of the machine learning model but are not trained during the machine learning training process. For example, a neural network designer decides the number of hidden layers and the number of nodes in each layer. Another example is that XGBoost is an open-source software library that implements machine learning algorithms under the gradient boosting framework and can include multiple hyperparameters, such as the number of trees, the maximum depth of the trees, the learning rate, the regularization parameter, and the number of different classes for classification problems. In various implementations, hyperparameters can be discrete and / or continuous and have a distribution of values described by hyperparameter expressions. The performance of a machine learning model depends largely on its hyperparameters. Summary of the Invention

[0002] In some aspects, the techniques described herein relate to a method for adjusting hyperparameters of a machine learning model. The method includes: for each hyperparameter, generating performance attribution statistics corresponding to the evaluation metric of the machine learning model based on historical experiment statistics for the evaluation metric and the machine learning model; assigning weights to each hyperparameter based on the performance attribution statistics of the hyperparameters; in a series of experiments, updating the hyperparameters based on the weights assigned to each hyperparameter; and selecting, from one of the experiments, a set of hyperparameters for the machine learning model, where the set of hyperparameters produces a recorded value of the evaluation metric that satisfies the adjustment condition.

[0003] In some aspects, the techniques described herein relate to a computing system for adjusting hyperparameters of a machine learning model. The computing system includes: one or more hardware processors; a performance attributor executable by the one or more hardware processors and configured to generate, for each hyperparameter, performance attribution statistics corresponding to the evaluation metric of the machine learning model based on historical experiment statistics for the evaluation metric and the machine learning model; a hyperparameter weight evaluator executable by the one or more hardware processors and configured to assign weights to each hyperparameter based on the performance attribution statistics of the hyperparameters; a hyperparameter updater executable by the one or more hardware processors and configured to update the hyperparameters based on the weights assigned to each hyperparameter in a series of experiments; and a hyperparameter selector executable by the one or more hardware processors and configured to select, from one of the experiments, a set of hyperparameters for the machine learning model, where the set of hyperparameters produces a recorded value of the evaluation metric that satisfies the adjustment condition.

[0004] In some aspects, the techniques described herein relate to one or more tangible processor-readable storage media that carry instructions for performing a process of tuning hyperparameters of a machine learning model on one or more processors and circuitry of a computing device, the process including: for each hyperparameter, generating performance attribution statistics corresponding to an evaluation metric of the machine learning model based on historical experiment statistics for the evaluation metric and the machine learning model; assigning a weight to each hyperparameter based on the performance attribution statistics of the hyperparameter; in a series of experiments, updating the hyperparameters based on the weights assigned to each hyperparameter; and selecting, from one of the experiments, a set of hyperparameters for the machine learning model, wherein the set of hyperparameters yields a recorded value of the evaluation metric that satisfies a tuning condition.

[0005] This Summary is provided to introduce a series of concepts in a simplified form that will be further described in the Detailed Description below. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.

[0006] Other implementations are also described and recited herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] Figure 1 Illustrates an example system for tuning hyperparameters of a machine learning model.

[0008] Figure 2 Illustrates an example hyperparameter tuner system.

[0009] Figure 3 Illustrates an example operation for tuning hyperparameters of a machine learning model.

[0010] Figure 4 Illustrates an example operation for updating hyperparameters within a pre-specified computational budget.

[0011] Figure 5 Illustrates an example computing device for use with deferred formula calculations. DETAILED DESCRIPTION

[0012] Hyperparameter tuning is the process of determining a hyperparameter configuration to yield a desired level of performance (e.g., optimal performance). Different types of performance can be measured by evaluation metrics such as one or more of accuracy, recall, specificity, sensitivity, F1 score, AUC-ROC, log loss, etc. This process can be computationally expensive and / or require manual management in some implementations because it may involve exploring a large range of values defined for each hyperparameter (e.g., grid search method) and manual selection of subsets of hyperparameters (e.g., random selection / expert selection). Such an example tuning process typically does not coordinate with a computational budget defined for the tuning objective (e.g., the amount of computational resources and / or number of computational cycles allocated for tuning, such as the number of allocated experiments).

[0013] In contrast, the described techniques are capable of providing technical benefits of reducing computational overhead and the need for manual intervention while coordinating a specified computational budget to quickly arrive at a set of tuned hyperparameters. In one implementation, by using a database of historical experiment statistics for hyperparameters of a particular machine learning model type and applying growth rate related criteria. For example, for these statistics, the described techniques can identify an initial value for each hyperparameter, determine the relative impact of each individual hyperparameter on the performance of the model, and apply appropriate weights to them. In this way, the allocation of the computational budget focuses on the higher-weight hyperparameters to obtain a set of tuned hyperparameters.

[0014] Designing a machine learning model typically involves running multiple experiments during the model development process and obtaining different results from the machine learning model. The experiment tracking database includes historical experiment statistics related to these previously executed experiments, such as how the evaluation metrics change with respect to different hyperparameters. Example experiments can evaluate but are not limited to different machine learning models, different model architectures, different hyperparameters, different training data, different evaluation metrics, different program codes, and / or the same program code running in different environments.

[0015] Figure 1 An example system 100 for tuning hyperparameters of a machine learning model is illustrated. Hyperparameters define the structural and algorithmic characteristics of a machine learning model but are not trained during the machine learning training process. Designing a machine learning model typically involves running different versions of the machine learning model (e.g., versions with different hyperparameters) with reference to one or more evaluation metrics. In this way, historical values of the (multiple) evaluation metrics are tracked for different values of each hyperparameter in a set of historical experiment statistics.

[0016] In various implementations, the hyperparameter tuner 102 receives a set of historical experiment statistics 104, which may be provided by an experiment tracking system that records machine learning experiments previously performed on a machine learning model for a specified evaluation metric. The historical experiment statistics 104 may include data characterizing the sensitivity of the evaluation metric to changes in the various hyperparameters of the model. The hyperparameter tuner 102 also receives a computational budget 106, which defines or may be translated into the number of experiments that can be allocated to a given tuning session. For example, the computational budget 106 may allocate one hundred experiments for tuning the hyperparameters of a machine learning model.

[0017] In at least one implementation, the hyperparameter tuner 102 also receives an untuned machine learning model 108. For example, the hyperparameters of the untuned machine learning model 108 may be initialized with associated but untuned hyperparameter values, which will be updated by the hyperparameter tuner 102 to tuned hyperparameter values. An experiment tracking database may be used to generate an initial version of the hyperparameters by setting each hyperparameter to the average of those values known to have resulted in the best value of the evaluation metric for that particular target machine learning model or model type in previous experiments, although other initialization operations may also be employed. The purpose of this initialization step is to start the hyperparameter search from a good initial point, thus accelerating the overall search time. The hyperparameter tuner 102 also receives training data 110, which may be used in the update operations of the hyperparameter tuner 102 during tuning.

[0018] In the illustrated implementation, the hyperparameter tuner 102 includes one or more components that, for each hyperparameter, generate performance attribution statistics corresponding to the evaluation metric of the machine learning model based on the historical experiment statistics for the evaluation metric and the machine learning model. Other components assign weights to each hyperparameter based on the performance attribution statistics of the hyperparameters, update the hyperparameters based on the weights assigned to each hyperparameter in a series of experiments, and select a set of hyperparameters for the machine learning model from one of these experiments. The selected set of hyperparameters produces recorded values of the evaluation metric that most satisfy the tuning criteria, and this set of hyperparameters is output as the tuned hyperparameters 112, which can then be used to design a tuned machine learning model 114.

[0019] Hyperparameters can be adjusted for different evaluation metrics. For example, the adjustment condition can be configured to determine whether the evaluation metric of "accuracy" of a machine learning model designed with a specific set of hyperparameters obtains the highest accuracy in the inference results of the machine learning model. Other adjustment conditions can be configured to determine whether the evaluation metric of "specificity" of a machine learning model designed with a specific set of hyperparameters obtains the highest true negative ratio correctly predicted by the machine learning model, or whether the evaluation metric of "sensitivity" of a machine learning model designed with a specific set of hyperparameters obtains the highest true positive ratio correctly predicted by the machine learning model. The described techniques can be used to adjust hyperparameters for other evaluation metrics.

[0020] The machine learning model can be trained for a specific task, such as identifying the category of an object shown in an image, or performing speech recognition on a voice signal. The training can be supervised training using labeled training data, such as training by a backpropagation training algorithm, or unsupervised learning can also be used. Any suitable training data can be used, such as publicly available datasets like ImageNet or other datasets.

[0021] In some cases, the machine learning model is used to control an autonomous vehicle or a communication network. Sensor data collected from the autonomous vehicle or the communication network is processed by the machine learning model to predict the values of one or more control parameters collected by the autonomous vehicle or the communication network. Then, these control parameters are automatically updated using the prediction results from the machine learning model. In some cases, the machine learning model is a generative machine learning model for generating speech or images.

[0022] Figure 2 An example hyperparameter tuner system 200 is illustrated. The hyperparameter tuner 202 includes various components configured to adjust hyperparameters for a specified evaluation metric. The evaluation metric corresponds to the performance objective of the machine learning model for which the hyperparameters are being adjusted.

[0023] The communication interface 204 receives inputs such as Figure 1 historical experiment statistics 104, a computational budget 106, an unadjusted machine learning model 108, and training data 110, and passes the inputs to the hyperparameter tuner 202. The communication interface 204 can include software and / or circuitry and can be executed by one or more hardware processors of the hyperparameter tuner system 200.

[0024] The performance attributor 206 is executable by one or more hardware processors of the hyperparameter tuner system 200 and is configured to generate performance attribution statistics for each hyperparameter. Each performance attribution statistic corresponds to an evaluation metric of a machine learning model, which is based on historical experiment statistics for the evaluation metric and the machine learning model, and the performance attribution statistic corresponding to a hyperparameter indicates the sensitivity of the evaluation metric to changes in the hyperparameter.

[0025] The hyperparameter weight evaluator 208 is executable by one or more hardware processors and is configured to assign a weight to each hyperparameter based on the performance attribution statistics of the hyperparameter. The weight (w i ) of each hyperparameter is used to influence the selection of the hyperparameter to be updated at each iteration. In one implementation, the hyperparameter to be updated at each iteration is randomly selected with probability w i . Accordingly, a hyperparameter with a higher weight has a higher probability of being selected for update in any particular iteration, thus tending to receive a larger number of update iterations on the more heavily weighted hyperparameters. Since it is known that the evaluation metric (based on historical experiment statistics) is more sensitive to changes in the more heavily weighted hyperparameters, the hyperparameter tuner system 200 focuses on updating those more heavily weighted hyperparameters as compared to the less heavily weighted hyperparameters, thereby making better use of the available computational budget.

[0026] The hyperparameter search initializer 210 initializes the weights of the hyperparameters in an un-tuned version of the machine learning model. An experiment tracking database can be used to generate an initial version of the hyperparameters by setting each hyperparameter to the average of those values known to have produced the best values of the evaluation metric in previous experiments for the machine learning model, although other initialization operations can also be employed. The purpose of this initialization step is to start the hyperparameter search from good initial values, thereby accelerating the overall search time. In other implementations, the un-tuned version of the machine learning model can be initialized before being input to the hyperparameter tuner 202.

[0027] The hyperparameter updater 212 is executable by one or more hardware processors and is configured to update the hyperparameters based on the weights assigned to each hyperparameter in a series of experiments. In one implementation, the hyperparameter updater 212 performs updates for a number of iterations (e.g., one experiment per iteration) limited by the computational budget by selecting a hyperparameter based on the weight assigned to the hyperparameter, updating the hyperparameter to a new value, performing an experiment on the machine learning model based on the new value of the hyperparameter, and recording the value of the evaluation metric resulting from the experiment. Various update methods can be employed. In one implementation, using a Bayesian update method, the hyperparameter updater 212 updates the hyperparameter to a new value based on maximizing the growth rate of the evaluation metric based on the change in the hyperparameter and minimizing the covariance of the evaluation metric based on the change in the hyperparameter.

[0028] The hyperparameter selector 214 is executable by one or more hardware processors and is configured to select, from among experiments in an experiment, a set of hyperparameters for a machine learning model, where the set of hyperparameters yields a recorded value of an evaluation metric that satisfies an adjustment condition.

[0029] The selected set of hyperparameters is output as adjusted hyperparameters 216 via a communication interface 218 (which may be the same interface as communication interface 204). The adjusted hyperparameters 216 can then be used to complete the design of the adjusted machine learning model. For example, if one of the adjusted hyperparameters is the number of layers of a neural network, the adjusted machine learning model is generated with the adjusted number of layers.

[0030] Figure 3 An example operation 300 for adjusting hyperparameters of a machine learning model is illustrated. A generation operation 302 generates performance attribution statistics for each hyperparameter related to an evaluation metric (e.g., performance, classification accuracy, log loss, confusion matrix) of the machine learning model. The performance attribution statistics corresponding to a hyperparameter indicate the sensitivity of the evaluation metric to changes in the hyperparameter. In some implementations, performance metric statistics from historical experiments are tracked in an experiment tracking database. These statistics are filtered down to include those records related to a specific machine learning algorithm and an evaluation metric for which hyperparameters are being adjusted. For example, the tracked statistics can be filtered to include only those statistics related to accuracy.

[0031] The terms ML_model and evaluationMetric refer, respectively, to the machine learning model and the performance objective for which hyperparameters are being adjusted. By fixing ML_model and evaluationMetric to selected values, the filtered statistics can be stored in a file with an example schema as follows, although other schemas may also be employed:

[0032] h1,..., h Nh , valueOfEvaluationMetric

[0033] where Nh represents the number of hyperparameters included in the experiment.

[0034] The rows in the file correspond to previously executed experiments, based on historical values of all hyperparameters and values of the evaluation metric that the target machine learning model attempts to optimize. The file defines a performance attribution statistical dataset. Subsequent steps include building a predictive series machine learning model whose task is to use the values of the hyperparameters as features to predict the value of the evaluation metric. The performance attribution dataset can be partitioned into a training subset and a test subset. Once the machine learning model is trained using the training subset, the test subset is applied to the machine learning model along with a local interpretability framework (such as LIME or SHAP) to give each hyperparameter a score indicating how much each hyperparameter contributes to the prediction of the performance attribution model. Hyperparameters with positive scores can be considered "high performers", while those with negative scores tend to drag down the performance. Further, hyperparameters with scores close to zero do not affect the overall performance in any significant way. Finally, scores are obtained for each hyperparameter, as shown in the following example scenario:

[0035] (First test sample) h 1_1_score , …, h 1_Nh_score

[0036] …

[0037] (Last test sample) h Ntest_1_score , …, h Ntest_Nh_score

[0038] where Ntest represents the number of experiments recorded in the file.

[0039] This data is used to calculate the following historical performance for each hyperparameter:

[0040] Average Score

[0041] m l = (h l_l_score +... + h Ntest_l_score ) / Ntest

[0042] …

[0043] m Nh = (h l_Nh_score +... + h Ntest_Nh_score ) / Ntest

[0044] Covariance Coefficient between all Hyperparameter Pairs (i, j)

[0045] sigma ij = ([h l_l_score - m i )(h l_j - m j ) +... + (m Ntest_i - m i)(m Ntest_j -m j )} / (2·Ntest 2 )

[0046] The allocation operation 304 assigns weights to each hyperparameter based on the performance attribution statistics of the hyperparameters. In various implementations, one or more sensitivity criteria are applied to each hyperparameter to assign a weight to each hyperparameter, and the weight includes the sensitivity of the evaluation metric to the change of the hyperparameter. Such criteria can be used to maximize the expected growth rate and median final value of the evaluation metric of the machine learning model. In one example, the Kelly criterion (e.g., generating a Kelly score or weight) is applied to determine such a weight, but other criteria can also be applied. The weights are represented herein as W = {w1,..., w Nh}, the set of weights assigned to each parameter, where the sum of the weights is normalized to 1 in at least one implementation.

[0047] In one implementation, the following example criterion objective ("Obj") is adopted:

[0048] Obj = [∑·D{-·∑[D·D·{{[[,

[0049] such that w1+..., w Nh = 1. The notation represents the search for the W value such that the expression within the square brackets is maximized:

[0050] 1. The first term within the square brackets attempts to maximize the total expected improvement in the evaluation metric - multiplying the average performance m i of the hyperparameter by its assigned weight w i and summing this product over all hyperparameters.

[0051] 2. The second term attempts to minimize the total covariance, which quantifies the level of risk / volatility of the evaluation metric - multiplying each covariance component of hyperparameters i and j by their respective assigned weights w i and w j and summing this product over all hyperparameter pairs. The sum of the assigned weights equals 1.

[0052] Generally, a numerical solver for (constrained) quadratic optimization can be used to solve the criterion objective problem. In some implementations, the criterion objective problem can be accurately solved in the absence of correlation between any hyperparameters, such that the optimal weights are given by:

[0053]

[0054] where is the standard deviation of the score. The result of the quadratic solver includes the set of weights W = {w1,..., w associated with Nh hyperparametersNh}.

[0055] Update operation 306 updates each hyperparameter using a hyperparameter tuning model based on the weights assigned to the hyperparameters. Figure 4 An example hyperparameter update process is described in more detail. A selection operation 308 selects a set of hyperparameters for the machine learning model from one of the experiments, where the set of hyperparameters produces recorded values of the evaluation metric that satisfy the adjustment condition.

[0056] Figure 4 An example operation 400 for updating hyperparameters within a pre-specified computational budget is illustrated. An initial version of the Nh hyperparameters is generated using the experiment tracking database by setting each hyperparameter to the average of those values known to have produced the best value of the evaluation metric (for that particular target machine learning model). The purpose of this initialization step is to start the hyperparameter search from a good starting point, thereby speeding up the overall search time. budget to set the number of experiments available for this hyperparameter tuning session, so this initialization can reduce the number of experiments required to obtain effective tuning.

[0057] For each experiment, the selection operation 402 is used to weight the obtained probability w of the distribution process. i Randomly choose the hyperparameter N i In this regard, those hyperparameters to which the evaluation metric is most sensitive (i.e., those with higher weights) have a higher probability of being selected and updated in each experiment. Accordingly, higher-weighted hyperparameters will typically be updated more times than lower-weighted hyperparameters.

[0058] Thereafter, an update operation 404 uses a dedicated Bayesian model b for each hyperparameter, such as i Update the model to update each hyperparameter h i In one example of updating the model, the Bayesian theorem can be used to update the hyperparameters. Note that in at least some implementations, each hyperparameter is assigned its own Bayesian model, so the database B = {b1, ..., b Nh}Model.

[0059] In general, the concept of Bayesian updating is to make informed decisions about the choice of hyperparameters to be used in an experiment taking into account past experience. In one implementation, Bayesian updating is used as an efficient method to converge to an adjusted (e.g., optimal) value of an evaluation metric by changing the value of the hyperparameter in different experiments (e.g., iterations). Each iteration of Bayesian updating provides a new value of the hyperparameter to test its effect on the evaluation metric.

[0060] Bayes' theorem is given by

[0061]

[0062] where is the hypothesized probability, which is the "prior" - the probability of being correct assuming no evidence; is the likelihood, which is the probability of the known evidence being correct given the hypothesis; is the probability of the evidence being correct given the hypothesis; and is the probability of the evidence, which is the sum of the products of the likelihood and the prior:

[0063] [D |

[0064] Accordingly, Bayesian updating can be used to update a hypothesis (e.g., the value of a hyperparameter to be used in a next experiment) when new data becomes available (e.g., the value of an evaluation metric from a previous experiment). For example, given data D as evidence such that the data point d i ∈ D is the value of an evaluation metric from experiment i, the posterior is:

[0065]

[0066] Given another data point d2 (e.g., the value of an evaluation metric from a second experiment), the posterior can be updated in a subsequent iteration, where the prior for that iteration is the posterior from the previous iteration:

[0067] |1

[0068] This way of updating can be propagated through multiple iterations of updates within the computational budget.

[0069] The experimental operation 406 performs an experiment on the machine learning model based on the updated value of the hyperparameter. For example, the experimental operation 406 executes the machine learning model on the training data set using the updated hyperparameters, thereby obtaining an evaluation metric for that experimental iteration (e.g., the obtained accuracy metric, the obtained specificity metric).

[0070] The recording operation 408 records the value of the evaluation metric obtained from the experiment associated with the set of hyperparameters used in the corresponding experiment. For example, after each experiment is completed, the results of the experiment are stored in a file with a schema such as the following:

[0071] h 1_(trial=1) , …, h Nh (trial=1) , valueOfEvaluatonMetric

[0072] …

[0073] h 1_(trial=t) , …, h Nh (trial=1) , valueOfEvaluationMetric

[0074] The decision operation 410 determines whether the number of experiments in the tuning session meets the computing budget. If not, another experiment is started via the selection operation 402 to continue the series of experiments. In the alternative, after all experiments have been performed (e.g., the computing budget has been exhausted), another selection operation 412 selects the hyperparameters associated with the best valueOfEvaluationMetric (e.g., the value that best meets the tuning criteria) and assigns them to the machine learning model to produce a tuned machine learning model.

[0075] Figure 5 An example computing device 500 for use with a deferred formula is illustrated. The computing device 500 can be a client device such as a laptop computer, mobile device, desktop computer, tablet computer, or server / cloud device. The computing device 500 includes one or more processors 502 and a memory 504. The memory 504 generally includes volatile memory (e.g., RAM) and non-volatile memory (e.g., flash memory). An operating system 510 resides in the memory 504 and is executed by the processor(s) 502.

[0076] In the example computing device 500, as Figure 5 shown, one or more modules or segments such as an application 550, a communication interface, a performance attributor, a hyperparameter weight evaluator, a hyperparameter search initializer, a hyperparameter updater, a hyperparameter selector, and other program code and modules are loaded into the operating system 510 on the memory 504 and / or a storage device 520 and executed by the processor(s) 502. The storage device 520 can store historical experiment statistical training data, computing budgets, evaluation metrics, hyperparameters, and other data and can be local to the computing device 500 or can be remote and communicatively connected to the computing device 500. In particular, in one implementation, the components of the hyperparameter tuner system can be implemented entirely in hardware or in a combination of hardware circuitry and software.

[0077] The computing device 500 includes a power supply 516 that is powered by one or more batteries or other power sources and provides power to the other components of the computing device 500. The power supply 516 can also be connected to an external power source that overrides or recharges the built-in battery or other power source.

[0078] The computing device 500 can include one or more communication transceivers 530 that can be connected to one or more antennas 532 to provide a network connection (such as a mobile phone network, ) The computing device 500 may also include a communication interface 536 (such as a network adapter or I / O port as a type of communication device). The computing device 500 may use the adapter and any other type of communication device to establish connections over a wide area network (WAN) or a local area network (LAN). It should be understood that the network connections shown are exemplary and other communication devices and components for establishing communication links between the computing device 500 and other devices may be used.

[0079] The computing device 500 may include one or more input devices 534 such that a user can input commands and information (e.g., a keyboard or a mouse). These and other input devices may be coupled to the server through one or more interfaces 538 (such as a serial port interface, a parallel port, or a universal serial bus (USB)). The computing device 500 may also include a display 522, such as a touch screen display.

[0080] The computing device 500 may include various tangible processor-readable storage media and intangible processor-readable communication signals. Tangible processor-readable storage may be implemented by any available medium accessible to the computing device 500 and may include volatile and non-volatile storage media as well as removable and non-removable storage media. Tangible processor-readable storage media do not include intangible communication signals (such as signals themselves) and include volatile and non-volatile, removable and non-removable storage media implemented in any method or technology for storing information such as processor-readable instructions, data structures, program modules, or other data. Tangible processor-readable storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic tape cartridges, tapes, magnetic disk storage or other magnetic storage devices, or any other tangible medium that can be used to store the desired information and can be accessed by the computing device 500. In contrast to tangible processor-readable storage media, intangible processor-readable communication signals may embody processor-readable instructions, data structures, program modules, or other data residing in a modulated data signal such as a carrier wave or other signal transmission mechanism. The term "modulated data signal" means a signal that has one or more of its characteristics set or changed in such a manner as to encode information in the signal. By way of example and not limitation, intangible communication signals include signals propagated through wired media such as a wired network or a direct line connection and wireless media such as acoustic, RF, infrared, and other wireless media.

[0081] Clause 1. A method for adjusting hyperparameters of a machine learning model, the method comprising: receiving an image or a voice signal; for each hyperparameter, generating performance attribution statistics corresponding to an evaluation metric of the machine learning model based on historical experiment statistics for the evaluation metric and the machine learning model; assigning a weight to each hyperparameter based on the performance attribution statistics of the hyperparameter; in a series of experiments, updating the hyperparameters based on the weights assigned to each hyperparameter; and selecting, from one of the experiments, a set of hyperparameters for the machine learning model, wherein the set of hyperparameters produces a recorded value of the evaluation metric that satisfies an adjustment condition; and using the machine learning model configured with the selected set of hyperparameters to perform speech recognition on the voice signal or object recognition on the image.

[0082] Clause 2. The method according to Clause 1, wherein the evaluation metric corresponds to a performance target of the machine learning model to which the hyperparameters are adjusted.

[0083] Clause 3. The method according to Clause 1, wherein the historical experiment statistics track historical values of the evaluation metric for different values of each hyperparameter.

[0084] Clause 4. The method according to Clause 1, wherein the performance attribution statistics corresponding to a hyperparameter indicate the sensitivity of the evaluation metric to changes in the hyperparameter.

[0085] Clause 5. The method according to Clause 1, wherein the update comprises: for multiple iterations limited by a computational budget, selecting a hyperparameter based on the weight assigned to the hyperparameter, updating the hyperparameter to a new value, performing an experiment on the machine learning model based on the new value of the hyperparameter, and recording the value of the evaluation metric resulting from the experiment.

[0086] Clause 6. The method according to Clause 1, wherein the update comprises: updating the hyperparameter to a new value based on maximizing the growth rate of the evaluation metric based on changes in the hyperparameter and minimizing the covariance of the evaluation metric based on changes in the hyperparameter.

[0087] Clause 7. The method according to Clause 1, wherein the update comprises: using a Bayesian model to update the hyperparameter to a new value.

[0088] Clause 8. A computing system includes: one or more hardware processors; a memory storing a machine learning model with hyperparameters, the machine learning model being configured to control an autonomous vehicle or a communication network; a performance attributor executable by the one or more hardware processors and configured to generate, for each hyperparameter, performance attribution statistics corresponding to an evaluation metric of the machine learning model based on historical experiment statistics for the evaluation metric and the machine learning model; a hyperparameter weight evaluator executable by the one or more hardware processors and configured to assign a weight to each hyperparameter based on the performance attribution statistics of the hyperparameter; a hyperparameter updater executable by the one or more hardware processors and configured to update the hyperparameters based on the weights assigned to each hyperparameter in a series of experiments; and a hyperparameter selector executable by the one or more hardware processors and configured to select, from one of the experiments, a set of hyperparameters for the machine learning model, wherein the set of hyperparameters yields a recorded value of the evaluation metric that meets an adjustment condition.

[0089] Clause 9. The computing system of Clause 8, wherein the evaluation metric corresponds to a performance target of the machine learning model to which the hyperparameters are adjusted.

[0090] Clause 10. The computing system of Clause 8, wherein the historical experiment statistics track historical values of the evaluation metric for different values of each hyperparameter.

[0091] Clause 11. The computing system of Clause 8, wherein the performance attribution statistics corresponding to a hyperparameter indicate the sensitivity of the evaluation metric to changes in the hyperparameter.

[0092] Clause 12. The computing system of Clause 8, wherein for multiple iterations limited by a computing budget, the hyperparameter updater is further configured to: randomly select a hyperparameter based on the weight assigned to the hyperparameter, update the hyperparameter to a new value, perform an experiment on the machine learning model based on the new value of the hyperparameter, and record the value of the evaluation metric resulting from the experiment.

[0093] Clause 13. The computing system of Clause 8, wherein the hyperparameter updater is further configured to update the hyperparameters to new values based on maximizing the growth rate of the evaluation metric based on changes in the hyperparameters and minimizing the covariance of the evaluation metric based on changes in the hyperparameters.

[0094] Clause 14. The computing system of Clause 8, wherein the hyperparameter updater is further configured to update the hyperparameters to new values using a Bayesian model.

[0095] Clause 15. One or more tangible processor-readable storage media carrying instructions for performing a speech generation process or an image generation process on one or more processors and circuitry of a computing device, the process comprising: for each hyperparameter, generating performance attribution statistics corresponding to an evaluation metric of a machine learning model based on historical experimental statistics for the evaluation metric and the machine learning model; assigning weights to each hyperparameter based on the performance attribution statistics of the hyperparameters; in a series of experiments, updating the hyperparameters based on the weights assigned to each hyperparameter; and selecting, from one of the experiments, a set of hyperparameters for the machine learning model, wherein the set of hyperparameters yields a recorded value of the evaluation metric that meets an adjustment condition; and using the machine learning model configured with the selected set of hyperparameters to perform speech generation or perform image generation.

[0096] Clause 16. The one or more tangible processor-readable storage media of clause 15, wherein the evaluation metric corresponds to a performance target of the machine learning model to which the hyperparameters are adjusted.

[0097] Clause 17. The one or more tangible processor-readable storage media of clause 15, wherein the historical experimental statistics track historical values of the evaluation metric for different values of each hyperparameter.

[0098] Clause 18. The one or more tangible processor-readable storage media of clause 15, wherein the performance attribution statistics corresponding to a hyperparameter indicate the sensitivity of the evaluation metric to changes in the hyperparameter.

[0099] Clause 19. The one or more tangible processor-readable storage media of clause 15, wherein the update comprises: for multiple iterations limited by a computing budget, selecting hyperparameters based on the weights assigned to the hyperparameters, updating the hyperparameters to new values, performing an experiment on the machine learning model based on the new values of the hyperparameters, and recording the value of the evaluation metric resulting from the experiment.

[0100] Clause 20. The one or more tangible processor-readable storage media of clause 15, wherein the update comprises: updating the hyperparameters to new values based on maximizing the growth rate of the evaluation metric based on changes in the hyperparameters and minimizing the covariance of the evaluation metric based on changes in the hyperparameters.

[0101] Clause 21. A system for tuning hyperparameters of a machine learning model, the system comprising: means for generating, for each hyperparameter, performance attribution statistics corresponding to an evaluation metric of the machine learning model based on historical experiment statistics for the evaluation metric and the machine learning model; means for assigning a weight to each hyperparameter based on the performance attribution statistics of the hyperparameter; means for updating the hyperparameters based on the weights assigned to each hyperparameter in a series of experiments; and means for selecting, from one of the experiments, a set of hyperparameters for the machine learning model, wherein the set of hyperparameters yields a recorded value of the evaluation metric that satisfies a tuning condition.

[0102] Clause 22. The system according to Clause 21, wherein the evaluation metric corresponds to a performance target of the machine learning model to which the hyperparameters are tuned.

[0103] Clause 23. The system according to Clause 21, wherein the historical experiment statistics track historical values of the evaluation metric for different values of each hyperparameter.

[0104] Clause 24. The system according to Clause 21, wherein the performance attribution statistics corresponding to a hyperparameter indicate the sensitivity of the evaluation metric to changes in the hyperparameter.

[0105] Clause 25. The system according to Clause 21, wherein the means for updating comprises: means for selecting, for multiple iterations limited by a computational budget, a hyperparameter based on the weight assigned to the hyperparameter; means for updating the hyperparameter to a new value; means for performing an experiment on the machine learning model based on the new value of the hyperparameter; and means for recording the value of the evaluation metric resulting from the experiment.

[0106] Clause 26. The system according to Clause 21, wherein the means for updating comprises: means for updating the hyperparameter to a new value based on maximizing the growth rate of the evaluation metric based on changes in the hyperparameter and minimizing the covariance of the evaluation metric based on changes in the hyperparameter.

[0107] Clause 27. The system according to Clause 21, wherein the update comprises: means for updating the hyperparameter to a new value using a Bayesian model.

[0108] Some implementations can include articles of manufacture. An article of manufacture can include a tangible storage medium storing logic. Examples of storage media can include one or more types of computer-readable storage media capable of storing electronic data, including volatile memory or non-volatile memory, removable or non-removable memory, erasable or non-erasable memory, writable or rewritable memory, etc. Examples of logic can include various software elements, such as software components, programs, applications, computer programs, applications programs, system programs, machine programs, operating system software, middleware, firmware, software modules, routines, subroutines, operational segments, methods, procedures, software interfaces, application program interfaces (APIs), instruction sets, computing code, computer code, code segments, computer code segments, words, values, symbols, or any combination thereof. In one implementation, for example, an article of manufacture can store executable computer program instructions that, when executed by a computer, cause the computer to perform the methods and / or operations according to the described embodiments. The executable computer program instructions can include any suitable type of code, such as source code, compiled code, interpreted code, executable code, static code, dynamic code, etc. The executable computer program instructions can be implemented according to a predefined computer language, manner, or syntax for instructing a computer to perform a certain operational segment. The instructions can be implemented using any suitable high-level, low-level, object-oriented, visual, compiled, and / or interpreted programming language.

[0109] The implementations described herein are implemented as logical steps in one or more computer systems. The logical operations can be implemented as (1) a sequence of processor-implemented steps executed in one or more computer systems and (2) interconnected machine or circuit modules within one or more computer systems. The implementation is optional depending on the performance requirements of the computer system utilized. Accordingly, the logical operations that make up the implementations described herein are variously referred to as operations, steps, objects, or modules. Further, it should be understood that the logical operations can be performed in any order unless explicitly stated otherwise or the claim language inherently requires a specific order.

[0110] The term "set of hyperparameters" refers to one or more values, excluding the empty set.

Claims

1. A method, comprising: Receiving an image or a voice signal; Adjusting hyperparameters of a machine learning model by: For each hyperparameter, generating performance attribution statistics corresponding to the evaluation metric of the machine learning model from historical experiment statistics for the evaluation metric and the machine learning model; Assigning a weight to each hyperparameter, the weight being related to the performance attribution statistics of the hyperparameter; In a series of experiments, updating the hyperparameters according to the weights assigned to each hyperparameter; And Selecting a set of hyperparameters for the machine learning model from one of the experiments, wherein the set of hyperparameters produces a recorded value of the evaluation metric that meets an adjustment condition; And Using the machine learning model configured with the selected set of hyperparameters to perform speech recognition on the voice signal or object recognition on the image.

2. The method according to claim 1, wherein the evaluation metric corresponds to a performance target of the machine learning model to which the hyperparameters are adjusted.

3. The method according to any one of claims 1 or 2, wherein the historical experiment statistics track historical values of the evaluation metric for different values of each hyperparameter.

4. The method according to any one of the preceding claims, wherein the performance attribution statistics corresponding to the hyperparameter indicate the sensitivity of the evaluation metric to changes in the hyperparameter.

5. The method according to any one of the preceding claims, wherein the update comprises: For multiple iterations limited by a computational budget, Selecting hyperparameters according to the weights assigned to the hyperparameters, Updating the hyperparameters to new values, Performing an experiment on the machine learning model using the new values of the hyperparameters, and Recording the value of the evaluation metric generated from the experiment.

6. The method according to any one of the preceding claims, wherein the update comprises: Updating the hyperparameters to new values by using changes in the hyperparameters to maximize the growth rate of the evaluation metric and using changes in the hyperparameters to minimize the covariance of the evaluation metric.

7. The method according to any one of the preceding claims, wherein the update comprises: Updating the hyperparameters to new values using a Bayesian model.

8. A computing system, comprising: One or more hardware processors; A memory storing a machine learning model having hyperparameters, the machine learning model being configured to control an autonomous vehicle or a communication network; A performance attributor executable by the one or more hardware processors and configured to: for each hyperparameter, generate performance attribution statistics corresponding to the evaluation metric of the machine learning model from historical experiment statistics for the evaluation metric and the machine learning model; A hyperparameter weight evaluator executable by the one or more hardware processors and configured to: assign a weight to each hyperparameter, the weight being related to the performance attribution statistics of the hyperparameter; A hyperparameter updater, executable by the one or more hardware processors and configured to: in a series of experiments, update the hyperparameters according to the weights assigned to each hyperparameter; and A hyperparameter selector, executable by the one or more hardware processors and configured to: select, from one of the experiments, a set of hyperparameters for the machine learning model, wherein the set of hyperparameters yields a recorded value of the evaluation metric that satisfies the tuning condition.

9. The computing system according to claim 8, wherein the evaluation metric corresponds to a performance goal of the machine learning model to which the hyperparameters are tuned.

10. The computing system according to any one of claims 8 or 9, wherein the historical experiment statistics track historical values of the evaluation metric for different values of each hyperparameter.

11. The computing system according to any one of claims 8 to 10, wherein the performance attribution statistics corresponding to the hyperparameters indicate the sensitivity of the evaluation metric to changes in the hyperparameters.

12. The computing system according to any one of claims 8 to 11, wherein for multiple iterations limited by a computing budget, the hyperparameter updater is further configured to: randomly select hyperparameters according to the weights assigned to the hyperparameters, update the hyperparameters to new values, perform an experiment on the machine learning model using the new values of the hyperparameters, and record the value of the evaluation metric resulting from the experiment.

13. The computing system according to any one of claims 8 to 12, wherein the hyperparameter updater is further configured to: update the hyperparameters to new values by maximizing the growth rate of the evaluation metric using changes in the hyperparameters and minimizing the covariance of the evaluation metric using changes in the hyperparameters.

14. The computing system according to any one of claims 8 to 13, wherein the hyperparameter updater is further configured to: update the hyperparameters to new values using a Bayesian model.

15. One or more tangible processor-readable storage media carrying instructions for performing a speech generation process or an image generation process on one or more processors and circuitry of a computing device, the process comprising: for each hyperparameter, generating performance attribution statistics corresponding to the evaluation metric of the machine learning model from historical experiment statistics for the evaluation metric and the machine learning model; assigning a weight to each hyperparameter, the weight being related to the performance attribution statistics of the hyperparameter; in a series of experiments, updating the hyperparameters according to the weights assigned to each hyperparameter; selecting, from one of the experiments, a set of hyperparameters for the machine learning model, wherein the set of hyperparameters yields a recorded value of the evaluation metric that satisfies the tuning condition; and using the machine learning model configured with the selected set of hyperparameters to perform speech generation or perform image generation.

16. One or more tangible processor-readable storage media according to claim 15, wherein the evaluation metric corresponds to a performance objective of the machine learning model to which the hyperparameter is tuned.

17. One or more tangible processor-readable storage media according to any one of claims 15 or 16, wherein the historical experiment statistics track historical values of the evaluation metric for different values of each hyperparameter.

18. One or more tangible processor-readable storage media according to any one of claims 15 to 17, wherein the performance attribution statistics corresponding to the hyperparameter indicate the sensitivity of the evaluation metric to changes in the hyperparameter.

19. One or more tangible processor-readable storage media according to any one of claims 15 to 18, wherein the update comprises: For multiple iterations limited by a computational budget, Select a hyperparameter according to the weight assigned to the hyperparameter, Update the hyperparameter to a new value, Perform an experiment on the machine learning model using the new value of the hyperparameter, and Record the value of the evaluation metric resulting from the experiment.

20. One or more tangible processor-readable storage media according to any one of claims 15 to 19, wherein the update comprises: Updating the hyperparameter to a new value by maximizing the growth rate of the evaluation metric using a change in the hyperparameter and minimizing the covariance of the evaluation metric using a change in the hyperparameter.