Efficient configuration selection for automated machine learning
By training and testing candidate machine learning configurations on the sampling of the data set, and using confidence interval pruning methods, the problem of finding suitable configurations in the prior art is solved, and a faster approximate optimal configuration choice is achieved.
Patent Information
- Application Number
- CN201980054858.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-08-23
- Filing Date
- 2019-06-28
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2039-06-28
AI Technical Summary
The prior art requires a lot of time to train and test when finding suitable machine learning configurations, especially when dealing with large-scale data sets, and trials of a single configuration can take hours or days.
By training and testing candidate machine learning configurations on the sampling of the dataset, pruning with confidence intervals, gradually shrinking the configuration set until an approximate optimal configuration is left.
Significantly reducing the time to select the appropriate machine learning configuration from the possible set of configurations, enabling the identification of an approximate optimal configuration at dozens or hundreds of times faster with no more than 1% accuracy loss.
Smart Images

Figure CN113168591B_ABST
Abstract
Description
Technical Field
[0001] The disclosed subject matter relates to the field of automated machine learning, i.e., automated training and testing of multiple machine learning configurations for the purpose of identifying optimal or near-optimal configurations. Background Art
[0002] The creation of a machine learning solution for a new prediction task or data set typically involves selecting a suitable machine learning model and / or learning algorithm from a plurality of possible models / algorithms (e.g., linear or logistic regression, support vector machines, decision trees and random forests, artificial neural networks), the setting of associated hyperparameters, and the selection of a variety of preprocessing and characterization methods for the data provided as input to the model / algorithm. The combination of data preprocessing / characterization, model / algorithm, and hyperparameter selection is also collectively referred to herein as a "machine learning configuration."
[0003] The performance of machine learning solutions for prediction tasks is highly dependent on the selected machine learning configuration. Therefore, data scientists often spend a lot of time training and testing many possible configurations and identifying the best one among them. This process may involve dozens or hundreds of trials. Although various tools have been developed to automate these trials, both manual and automatic methods are becoming more and more time-consuming due to the growing datasets. For large-scale datasets, trials of just a single configuration may take hours or days. Therefore, a more efficient method is needed to select the appropriate machine learning configuration from a set of possible configurations. Summary of the invention
[0004] An automated machine learning method is described herein that generally involves training and testing a set of candidate machine learning configurations (also referred to herein as a "candidate set") on a sampled (rather than entire) data set to iteratively identify an optimal configuration or a near-optimal configuration. The identified configuration is also referred to herein as a "near-optimal configuration". In various embodiments, after the selected configuration is trained and tested on the sampled data set, the associated training accuracy and test accuracy (or the training and test values of some other quality metric of the trained configuration) are used to estimate the confidence interval (i.e., confidence upper and lower bounds) of the true performance of the configuration if it is trained and tested on the entire data set. Comparisons between the estimated confidence limits associated with various configurations are used to gradually "prune" the candidate set by removing configurations with low performance. In addition, the confidence interval can be iteratively fine-tuned by gradually increasing the sample size to repeatedly train and test any given configuration. The iterative training, testing, and pruning process can continue until only one configuration remains in the candidate set; this remaining configuration constitutes a near-optimal configuration and can be trained on the entire data set to optimize its performance. In various embodiments, pruning and calculation of confidence intervals are performed in a manner that ensures that, under a specified minimum probability, the accuracy (or other quality metric) of the approximate optimal configuration is within a specified loss tolerance of the accuracy (or other quality metric) of the "true best" (i.e., optimal) configuration.
[0005] Beneficially, the progressive sampling and pruning strategies described herein enable identification of a near-optimal machine learning configuration in significantly less time than it would take to determine the true optimal configuration by exhaustively training and testing on the entire dataset. For example, in some embodiments, a near-optimal configuration may be identified tens or hundreds of times faster with no more than a 1% loss in accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0006] The foregoing will be more readily understood from the following detailed description of various embodiments, particularly in conjunction with the accompanying drawings.
[0007] Figure 1 is a schematic block diagram illustrating a system for efficiently and automatically determining a near-optimal machine learning configuration according to various embodiments.
[0008] Figure 2 is a graph showing learning curves for multiple example machine learning algorithms trained on example data sets.
[0009] Figures 3A-3C A sequence of estimated confidence intervals for a set of candidate machine learning configurations is shown, illustrating stepwise pruning in accordance with various embodiments.
[0010] Figure 4is a flow chart illustrating a method for efficiently and automatically determining a near-optimal machine learning configuration according to various embodiments.
[0011] Figure 5 It is shown that according to various embodiments Figure 4 Flowchart of the iterative approach for selecting new machine learning configurations.
[0012] Figure 6 According to various embodiments, it can be used to implement Figure 1 A block diagram of an example computing system of a system. DETAILED DESCRIPTION
[0013] Described herein are systems, methods, and computer program products (embodied in machine-readable media) for efficiently and automatically selecting a machine learning configuration from a set of candidate configurations by iteratively training and testing the candidate configurations using stepwise sampling, combined with stepwise, confidence interval-based pruning of the candidate set. In various embodiments, the machine learning configuration ultimately selected is a near-optimal configuration, i.e., a configuration that achieves the best performance or near-optimal performance (as defined in a probabilistic sense according to some specified criterion or criteria) compared to other configurations in the candidate set.
[0014] Figure 1 An example computing system 100 for efficiently and automatically determining approximately optimal machine learning configurations according to various embodiments is shown. The computing system 100 may be implemented with a suitable combination of hardware and / or software, and typically includes one or more appropriately configured or programmed hardware processors (such as a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), etc.). In various embodiments, the computing system 100 is, for example, Figure 6 The general-purpose computer shown, or a cluster of multiple such computers communicating with each other via a network. In short, the computer or cluster includes one or more CPUs and / or GPUs, and (volatile and / or non-volatile) memory, which stores data and program code for execution by the CPU and / or GPU. The computer may also include input / output devices (e.g., keyboard / mouse and screen display for user interaction) and / or a network interface (e.g., for connecting to the Internet).
[0015] The computing system 100 operates on a set of candidate machine learning configurations 102 and a data set 104 for training and testing the candidate configurations, and returns an approximate optimal configuration 106 as output. The data set 104 may be divided (e.g., randomly) into a training data set 107 for training the candidate configurations and a test data set 108 for validating the trained configurations (i.e., testing their performance). The set of candidate machine learning configurations 102 and the data set 104 may be stored on one or more machine-readable media (e.g., in one or more databases) that are part of a computer implementing the computing system 100 or accessible by the computing system 100 via a network. For example, in one non-limiting embodiment, the computing system 100 and the candidate set 102 may be set on a server computer or a group of server computers accessible by client computers via the Internet. The data set 104 for a particular machine learning task may be stored on a client computer and may be remotely accessed by a server computer of the computing system 100, or alternatively, uploaded to a server computer. After determining the approximately optimal configuration 106, an identifier of the configuration 106 (including the data transformation operations, the name of the machine learning model, and the associated parameter values) can be transmitted to the client computer, for example, via a web-based user interface or message, and / or stored in the computing system 100 for later use by the client computer. Computer code and data structures that implement the selected configuration can also be downloaded to the client device.
[0016] The system 100 may also take as input (e.g., provided by a client computer) a performance and / or time constraint 110 that quantifies how close the returned configuration 106 is to the true best configuration and / or specifies a time limit for termination of the iterative configuration selection process. For example, the performance criteria may include an accuracy loss tolerance and an associated minimum probability (typically taken as a value close to 1, such as at least 95% or even 99%) that the accuracy of the last remaining configuration in the candidate set 102 returned as the approximate best configuration 106 does not differ from the accuracy of the true best configuration by more than a specified accuracy loss tolerance. Alternatively, or in addition, loss tolerances for other quality metrics used include, for example, loss tolerances for mean squared loss, normalized discounted cumulative gain, and area under the curve (AUC). In embodiments where time constraints are imposed, the pruned candidate set may still include multiple candidate configurations upon termination. Since the performance of these configurations (e.g., as measured by confidence bounds on their accuracy) is generally higher and more tightly clustered (i.e., characterized by narrower distribution, higher mean) than the performance of the initial candidate set, any of the remaining candidate configurations can be used as the approximately optimal configuration. Alternatively, the final selection step can identify the configuration with the highest performance (e.g., highest confidence lower bound) among the remaining configurations in the set at the termination time.
[0017] Each machine learning configuration within set 102 can specify a mathematical model (e.g., an equation or algorithm) for predicting output data based on input data, in conjunction with a learning algorithm for setting adjustable parameters of the model to fit training data set 107. The model and / or learning algorithm can also include hyperparameters, which are fixed for a given configuration, but can vary between multiple potential configurations of the model and a given type of learning algorithm. In addition, each machine learning configuration can specify how to pre-process and / or characterize "raw" input data (e.g., which can include numbers, text, images, or audio data) to generate digital inputs (e.g., input vectors) on which the model can operate. Therefore, different configurations within set 102 typically differ in one or more of the type of prediction model, the learning algorithm, hyperparameters associated with the model or learning algorithm, and the calculation and selection of input features.
[0018] The machine learning configuration that forms the candidate set 102 (including the types of models and algorithms contained therein) generally depends on the specific machine learning task and its associated data type. For example, for tasks involving predicting dependent quantitative variables based on independent quantitative variables, the candidate set 102 may include one or more regression models of decision trees and / or candidate functional relationships between specified variables. As another example, for classification tasks, the models within the candidate set 102 may include, but are not limited to: a naive Bayes classifier, a decision tree or random forest, a logistic regression model, and / or one or more artificial neural networks (possibly with various network architectures). The machine learning configuration of the neural network model may in turn specify various associated learning algorithms (e.g., error back propagation, or reinforcement learning with various rewards), as well as hyperparameter differences (such as the number of layers within the network, or the step size used when adjusting the network weights during the learning process).
[0019] In various embodiments, the system 100 is used to select a machine learning configuration for a supervised learning task. In supervised learning, the data set 104 consists of paired input and output items to provide a direct method to measure the performance of the trained machine learning configuration. For example, in a classification task, the output item is a label, each label specifies the category to which the corresponding input item belongs. A suitable quality metric for a trained classifier model is its classification accuracy, for example, measured by the fraction of input items that are correctly classified (i.e., consistent with the output label). The classifier model can be trained to maximize the classification accuracy on the training data set 107, and its performance can then be evaluated based on the classification accuracy achieved on the test data set 108. In the case of predicting a dependent variable based on an independent variable, the output item is the true value of the dependent variable given the independent variable input. In this case, the prediction accuracy of the training model can be determined based on the difference between the true output value and the predicted output value (e.g., measured by the sum of squared differences).
[0020] In various example embodiments described herein, the accuracy of a trained machine learning configuration is used as a quality metric to quantify its performance. However, it should be understood that alternative quality metrics (e.g., mean squared loss, discounted cumulative gain, or AUC) may also be used. In addition, determining the approximate optimal configuration based on stepwise sampling and pruning is not limited to configurations for supervised learning tasks, but can be similarly applied to unsupervised learning or reinforcement learning situations, where suitable quality metrics known to those of ordinary skill in the art can be used to measure the performance of the configuration and calculate associated confidence intervals. For example, for unsupervised learning tasks, mutual information and average distance are suitable quality metrics. In reinforcement learning, average regret can be used as a quality metric.
[0021] Re-reference Figure 1 , the computing system 100 may include multiple processing components to select a near-optimal configuration 106 from the candidate set 102; these components may include a training and testing component 112, a data sampler 114, and a planning and pruning component 116. The individual components 112, 114, 116 may be implemented as separate software programs or modules, for example, executed by a shared hardware processor. Alternatively, different components of the components 112, 114, 116, or even different subcomponents thereof, may be implemented by separate hardware components. For example, while the planning and pruning component 116 may be implemented in software executed by a CPU, the training and testing component 112 may use an FPGA to execute some learning algorithms for the candidate set 102, and other components may use a GPU or a CPU. In addition, in some embodiments, while functionally belonging to the computing system 100, the sampler 114 may be executed on a separate computer that holds the data set 104.
[0022] The training and testing component 112 is configured to train the selected candidate configuration on the sampled training data set 118, which generally involves executing the learning algorithm of the selected configuration to adjust the parameters of the associated model. The training and testing component 112 is also configured to evaluate the performance of the trained configuration on both the sampled training data set 118 and the sampled test data set 119 to calculate the associated training and test accuracy (or other quality metrics) 120. The sampler 114 is configured to generate the sampled training and test data sets 118, 119 by sampling (e.g., randomly sampling) from the complete training data set 107 and the test data set 108, respectively, using a sample size 122 determined by the planning and pruning component 116 and transmitted to the data sampler 114 (e.g., via the training and testing component 112). The data sampler 114 and the training and testing component 112 can be readily implemented by a person of ordinary skill in the art without undue experimentation. For example, existing publicly available software tools that implement the training and testing component 112 or portions thereof are included in the open source Machine Learning Toolkit "TLC" (developed by Microsoft Corporation of Redmond, Washington) and "scikit-learn."
[0023] The planning and pruning component 116 is configured to control the iterative process of sampling, training, and testing machine learning configurations selected from the candidate set 102, and pruning the candidate set 102. Based on the training and testing accuracies 120 calculated by the training and testing component 112, the planning and pruning component 116 calculates and updates confidence intervals associated with the trained and tested configurations, and then prunes the candidate set 102 based on this and in combination with the performance and time constraints 110. In some embodiments, the planning and pruning component 116 removes from the candidate set 102 any configuration whose upper confidence bound exceeds the highest lower confidence bound (among all configurations) by no more than a tolerance for accuracy loss. The planning and pruning component 116 further selects the next configuration 124 to be trained and tested in each iteration, and determines the associated sample size 122. The selected configuration 124 and sample size 122 can be transmitted to the training and testing component 112. Various functions of the planning and pruning component will be referred to below. Figure 3A-5 Described in further detail.
[0024] Now turn to Figure 2, some observations and insights about progressive sampling are illustrated with this graph, which shows empirical learning curves 200, 202, 204, 206, 208 for five exemplary machine learning configurations trained on an example dataset of flight delay data. Each learning curve plots the test accuracy of the corresponding configuration (determined for a constant (full) test dataset) as a function of the sample size of the sampled training dataset on a logarithmic scale. For sufficiently large sample sizes (greater than about 2 million), the configuration with the highest test accuracy is "LightGBM" (curve 204). It can also be seen that the "optimal" sample size (i.e., the minimum sample size beyond which the test accuracy is not significantly improved) for minimizing the running time of the training configuration varies between different configurations, being approximately 16,000 (compared to 2 million) for all other shown configurations (curves 200, 202, 206, 208). Therefore, if the optimal sample size is known in advance, these other configurations can be tested using only 16,000 samples, further reducing the overall training time. However, in general, the optimal sample size is unknown at the beginning. Moreover, a seemingly natural approach is to gradually increase the sample size during iterative training and testing until a plateau is reached in the learning curve and then use this sample size as an estimate of the optimal sample size, but this approach is prone to error. For example, the learning curve 204 of LightGBM is relatively flat between 32,000 and 128,000 samples, but the test accuracy increases significantly with more than 128,000 samples.
[0025] A more powerful strategy is disclosed herein that involves estimating confidence intervals for the true test accuracy of a configuration, rather than using point estimates such as a "plateau estimate." As the sample size increases during repeated training of a given configuration, the confidence intervals shrink, thereby pruning away poorly performing configurations.
[0026] Figures 3A-3C Stepwise confidence interval-based pruning according to various embodiments is shown, with an example sequence of estimated confidence intervals for an example candidate set of initial five machine learning configurations labeled C1 to C5. For ease of illustration, pruning is performed with an accuracy loss tolerance of zero in the described example, which means that a configuration is removed from the set only if its upper bound is equal to or lower than the lower bound of the highest confidence interval of all configurations (such that there is no longer an overlapping range between the two configurations). However, more generally, the accuracy loss tolerance need not be zero, but can also be set to a small positive value that allows pruning even in the case of configurations that have some overlap (up to a non-zero accuracy loss tolerance) with the highest confidence configuration.
[0027] Figure 3A The upper and lower confidence bounds for all five configurations at once in the iterative training and testing process are shown when all configurations are trained and tested on the corresponding sample data sets. The best performing configuration at this time is configuration C1. It can be seen that the upper bound 300 of the confidence interval of configuration C5 is lower than the lower bound 302 of the confidence interval of configuration C1. Therefore, configuration C5 can be deleted from the candidate set. Figure 3B The remaining candidate set with updated confidence intervals after some iterations is shown. Now, configuration C2 has exceeded the performance of configuration C1 and has the highest associated lower bound 304, and configuration C3 has an upper bound 306 that has fallen below the lower bound 304. Therefore, configuration C3 is deleted at this stage. Figure 3C As shown, there are still some iterations later, and the upper bounds 308, 310 of configurations C1 and C4 are below the updated lower bound 304 of configuration C2. Therefore, configurations C1 and C4 can be pruned away, leaving only configuration C2 as the approximately optimal configuration in the candidate set.
[0028] refer to Figure 4 , a method 400 for efficiently and automatically determining a near-optimal machine learning configuration according to various example embodiments will now be described in more detail. At 402, method 402 takes as input an initial set C (corresponding to set 102) of |C|=n candidate configurations, a training data set and a test data set (107, 108), and a specified loss tolerance (e.g., accuracy loss tolerance) ∈. Method 400 involves an iterative process of training and testing (also referred to herein as "heuristics") selected candidate configurations (also referred to herein as "heuristic configurations") and pruning a set of candidate remaining configurations Ω based on these heuristics. At 404, the remaining configuration set Ω is initialized to C, and the heuristic configuration C is prob Initialized to a candidate configuration C selected (e.g., randomly) from C 1 , and assume the optimal configuration C i Initialize to the same candidate configuration C 1 In addition, the estimated confidence intervals for all configurations in set C can be initialized, for example, to the full possible range that the selected quality metric can assume (e.g., to a range of 0 to 1 for accuracy-based confidence intervals). At 406, for example, a sample size for configuration C is determined based on a predetermined sampling plan associated with the configuration. 1 The initial training sample size can also be used for C 1 The sampling plan of determines the initial test sample size at 406. After these initialization operations, the method 400 enters a loop in which the selected configurations are iteratively trained and tested, and the set Ω is iteratively pruned as long as the number of remaining configurations in the set Ω is greater than 1 (determined at 408).
[0029] In each cycle of the iterative process, at 410, the training data set and the test data set are sampled (eg, by the data sampler 114) based on the determined sample size. At 412, the tentative configuration C is trained on the sampled training data set. prob , and then evaluate the tentative configuration C on a sampled test dataset (or in some embodiments, on the entire test dataset) prob (e.g., by training and testing component 112). During the training and testing process, the proposed heuristic configuration C that represents the trained one will be evaluated on the sampled training dataset and the test dataset. prob For example, if prediction accuracy is used as the quality metric, then the training and test accuracies are calculated. At 414, based on the training and test accuracies (or training and test values of some other quality metric), optionally in combination with other parameters, the heuristic configuration C is updated (e.g., by the planning and pruning component 116). prob The associated estimated confidence interval. Confidence interval for the trial configuration C prob The configuration provides an estimated bound on the true performance of , i.e., the accuracy (or other quality metric) that the configuration would achieve if trained on the entire training data set 107 and tested on the entire test data set 108. The lower bound of the estimated confidence interval is typically lower than the test accuracy (or the test value of another quality metric), and the upper bound of the estimated confidence interval is typically greater than the training accuracy (or the training value of another quality metric).
[0030] At 416, the updated lower bound C prob l and the currently assumed optimal configuration C i′ The lower bound C i′ l is compared, and if C prob l>C i′ ·, l, then the optimal configuration C will be assumed i′ Update to Trial Configuration C prob (And accordingly, the lower bound C i′ .l updated to C prob ·l). Then, at 418, based on the lower bound C of the new assumed optimal configuration i′ Compare l (which is the configuration with the highest lower bound due to iteratively updating the assumed best configuration) with the upper bound Cu of other configurations C in the set Ω, and prune the remaining set of configurations Ω: remove from the set Ω those configurations whose upper bound exceeds the highest lower bound by no more than the loss tolerance ∈ (Cu-C i′ ≤∈). (As will be appreciated by those skilled in the art, in alternative embodiments, the Cu-C i′=∈, and only prune the upper bound beyond the highest lower bound less than the loss tolerance∈ (i.e., Cu-C i′ <∈). Upper bound Cu=C i′ It does not really matter whether the configuration of +∈ is preserved or pruned, that is, the two embodiments are equivalent in practical applications.)
[0031] After pruning (at 418), at 420, the next tentative configuration is selected from the remaining set of configurations Ω, and the associated sample sizes for sampling the training and test data sets are determined (e.g., by the planning and pruning component 116). Configuration selection can be based on confidence intervals associated with the configurations and / or the training time required to reduce the confidence intervals by training with increased sample sizes; see below for more information. Figure 5 An example embodiment is described in detail. The iterative process of training and testing the selected configuration on the sampled data set, updating the confidence interval, pruning the set of candidate remaining configurations Ω, and selecting a new tentative configuration (operations 410-420) can continue as long as more than one candidate configuration remains in the set Ω. Once only one configuration remains in the set Ω, that configuration is returned as the approximately optimal configuration (at 422). Alternatively, in some embodiments ( Figure 4 ), the iterative process can terminate when a specified time limit has been reached; then, among the configurations still remaining in the set Ω, one configuration can be selected (e.g., based on the highest associated upper or lower bound) and returned as the approximately optimal configuration.
[0032] In order for the method 400 to efficiently identify the best or at least approximately best performing configuration among the candidate configurations with high probability (as evaluated below based on the accuracy of the trained model), the confidence interval can be calculated in a manner that satisfies the following two criteria: the confidence interval for a given trained configuration contains the true test accuracy of the configuration with high probability, and the calculation of the confidence upper and lower bounds is not slower than the training of the configuration. In various embodiments, these criteria are met by the following: upper and lower bounds calculated based on the training accuracy and test accuracy of the trained configuration, and in combination with the size of the sampled training data set, the sampled test data set, and the entire test data set (i.e., the number of samples), and the number of configurations in the initial candidate set, and the probability that the accuracy of the approximately best configuration returned by the method 400 is within the loss tolerance ∈ of the accuracy of the best configuration.
[0033] To be more specific, The confidence interval [l,u] associated with a given trial of is associated with the true performance of the configuration, let D tr and D te denote the training data set 107 and the test data set 108 respectively, and let S tr and Ste Denote the sampled training data set 118 and the sampled test data set 119 of the given heuristic sampled data set 119 respectively. The corresponding number of samples is represented by |D tr |、|D te |、|S tr | and |S te | indicates. In addition, let H tr Indicates that when in the entire training data set D tr When training on configuration C (also referred to as “on dataset D” in this paper), tr The machine learning model (e.g., classifier) output by the learning algorithm under the configuration C” trained on the training set, and similarly Indicates that in the sampled training data set S tr In addition, let and Indicates that in the sampled training data set S tr The training accuracy and test accuracy of configuration C trained on , and let Indicates that in the entire training data set D tr The true test accuracy of configuration C when training on .
[0034] Assuming that the test accuracy of a configuration trained on D in dataset D is no worse than the test accuracy of the same configuration trained on a different dataset D′ in D (which reflects that the training process generally produces a trained model that fits the training data), it can be shown that with probability at least (where n is the initial set In the case of the number of configurations in ), such that:
[0035]
[0036] The confidence upper bound u thus calculated has an additive form consisting of the following three components: tr The training accuracy on the training sample size, the variation term S caused by tr , and because the entire test data size |D te |. Intuitively, u increases with the training accuracy This is because higher training accuracy indicates higher potential of the configuration’s learning ability. However, as the training sample size |S tr This potential decreases as | increases, because the more data used in training, the smaller the improvement that can be obtained by adding more training data. Finally, since the true test accuracy is the average over the entire test dataset D te , thus adding the variation caused by the random split of the dataset into training and test datasets. teThe larger the value, the smaller the change. Both of these change terms are affected by the confidence probability The influence of . A high confidence probability corresponds to a wider confidence interval and, therefore, to a larger value of u. In summary, the confidence upper bound u calculated according to the above formula is positively correlated with the training accuracy and the number of configurations n, and negatively correlated with the size of the training dataset and the entire test dataset. Note that the training accuracy (This is the most time-consuming step in the calculation of u, and the calculation time depends on the sample size S tr The only step that changes is that the training is no slower than training on a configuration with a sampled dataset. In fact, testing is often much more efficient than training on a dataset of the same size.
[0037] Now turning to the confidence lower bound, since training on the entire dataset produces better accuracy than training on the sampled training dataset, we can get the test accuracy of the configuration trained on the sampled training dataset by as the true test accuracy of the trained configuration However, if |D te |>>|S tr |, then used to calculate Testing the trained configuration on the entire test dataset is slower than training on the sampled training dataset. In order to calculate the confidence interval, according to various embodiments, the test data is also sampled, and the test accuracy on the sampled test dataset minus a certain variation term is used as a lower bound on the test accuracy on the entire test dataset. It can be shown that with probability at least In this case,
[0038]
[0039] Therefore, the confidence lower bound l calculated is the sample test data set S te middle The accuracy is minus the sample size S due to sampling the test data set te As the sample size S te increase, and The difference between becomes smaller, and the lower bound also increases. Higher confidence probability corresponds to a smaller l. In summary, l is positively correlated with the size of the sampled test dataset and the test accuracy in that sample, and negatively correlated with n.
[0040] According to the above two inequality relations, it can be concluded that the true accuracy of the trained configuration is At least The probability that is within the confidence interval [l,u], where the above expressions represent the lower and upper bounds of the confidence interval. It can also be shown that the use of these expressions for l and u Figure 4 The method 400 may return as an approximate optimal configuration C i′ The actual test accuracy of this configuration Test accuracy in real-world optimal configurations The accuracy loss tolerance ∈ is within the range The probability is at least 1-δ.
[0041] Therefore, in various embodiments, an accuracy loss tolerance ∈ and a (smaller) maximum permissible probability δ that the identified approximate optimal configuration deviates from the optimal configuration by more than ∈ are specified as inputs to the configuration selection method, and then the approximate optimal configuration is determined based on the confidence interval, which is calculated according to the above formula by progressively pruning overlapping configurations whose associated confidence intervals overlap with the current confidence interval with the highest lower bound. According to some embodiments, the loss tolerance can be set to zero so that the identified approximate optimal configuration is the true optimal configuration (or one of multiple configurations with equal optimal performance).
[0042] Note that method 400 does not necessarily use the above specific expressions for upper and lower bounds. Instead, alternative formulas may be used to estimate confidence intervals, but in other cases, the above-mentioned probability guarantees of finding a configuration within a specified loss tolerance may not apply. Nevertheless, a pruning process that employs different estimates of confidence limits can provide an efficient way to determine a possible at least approximately optimal configuration or a reduced set of candidates for such a configuration. Subsequent training and testing can be used to further evaluate the performance of such a configuration.
[0043] Turning now to the selection of the tentative configuration and associated sample size (at 420 of iterative method 400), Figure 5 A method 500 is shown for selecting the next tentative configuration based on a training cost gradient that measures the increase in training time required to achieve a narrower confidence interval. To motivate the depicted method, consider the total runtime for identifying and training a near-optimal configuration Let T i (s) represents the configuration C on the sampled training dataset of size s i The trial time, and let t i is the cumulative running time of the trial configuration in method 400. In addition, let l i and u i C is the time when the iterative algorithm terminates i Without loss of generality, assume that method 400 returns C 1as the best approximation configuration. Using these notations, the total running time can be expressed as:
[0044]
[0045] According to various embodiments, the planning of the tentative configurations (i.e., the selection of the next tentative configuration in each configuration) is performed so that Designed with minimization as the goal, the constraints are as follows:
[0046] u 2 ≤l 1 +∈,u 3 ≤l 1 +∈,...,u n ≤l 1 +∈,
[0047] This ensures that C 1 All configurations except In the expression for , the first term corresponds to the time taken to identify the approximately optimal configuration; since the running time of each iteration is usually determined by training, the total time of all trials is used as a proxy for this identification time. The second term is the time taken to determine the approximately optimal configuration (once it has been identified) for training on the entire training dataset; this term is constant.
[0048] In order to Minimize, first, when the "oracle" optimal plan has access to C as a function of the upper and lower bounds of the corresponding confidence intervals at termination i The cumulative running time t i (i.e., t i =f i (l i ) = g i (u i )), the "oracle" optimal plan is studied. Accessed through this oracle, the optimal solution will only try each configuration once (because otherwise the total running time can be reduced by keeping only the last trial). Ignoring the constant term, the total running time can be Rewrite as f 1 (l 1 )+g 2 (u 2 )+…+g n (u n ). Using the Lagrange multiplier method, it can be shown that the optimal solution will satisfy:
[0049]
[0050] Among them l 1 +∈=u 2=…=u n .
[0051] In fact, there is no i (l i ) and g i (u i ) of Oracle access, and there is no closed-form formula to determine the configuration C i The optimal sample size Therefore, accordingly, the configuration is iteratively trained on a dataset with gradually increasing sample sizes. In some embodiments, the associated plan is informed by the above-mentioned oracle-based solution, and the associated plan uses the f i and g i Approximate the gradient of the training cost with respect to the derivative with respect to the confidence interval bounds.
[0052] refer to Figure 5 According to various example embodiments, at 502, the selection of the next configuration to be tried in the iteration process (at 420 in method 400) takes as input the current set Ω of m remaining configurations and their associated confidence intervals. The configurations are sorted (in descending order) according to their respective upper bounds, and the configuration with the highest upper bound (here, without loss of generality, assumed to be C) is selected. 1 ) as an initial guess of the best configuration (at 504). In addition, at 506, based on the corresponding configuration C i The running time difference ΔT between the two most recent consecutive trials i and the difference Δl between the associated confidence limits i and Δu i , calculate the training cost gradient for the m remaining configurations in Ω (for i=1) and (for i=2, ..., m). At 508, the training cost gradient of the configuration with the highest upper bound is is compared to the sum of the training cost gradients for all other configurations. If Then select configuration C 1 for the next trial (at 510); otherwise, the configuration with the second highest upper bound is selected for the next trial (at 512). Intuitively, if C 1 The lower bound of (the training time spent each time) increases faster than the upper bound of all other configuration combinations decreases, then the next trial C 1 Otherwise, choose the configuration with the second highest bound (in C 2 to C m), which will achieve the same upper bound for all configurations (as suggested by the second condition of the oracle-based solution). Once the configuration for the next trial is selected, the associated sample size for the next trial is determined (at 514), for example, based on the sampling plan associated with the selected configuration. The selected configuration and the associated sample size are then output (at 516) back to the iterative method 400.
[0053] According to various embodiments, the sample plan associated with a configuration is predetermined so that the sample size for the next trial (i.e., the size of the sampled training data set, and if the test data set is sampled in the same way, the sample size of the sampled test data set) can be simply found in an iterative process. The sample plan can be geometric, meaning that between any two consecutive trials of the same configuration, the sample size (either for training or testing) will increase by a constant factor c. The optimal value of c typically depends on some aspect of the configuration (e.g., the model and / or the learning algorithm). It can be shown that when the configuration C i The trial time T i (s) is the power function of the training sample size s, that is, T i (s)=s α (where α is a real number), the optimal step size meets For example, if the time used to explore a configuration is proportional to the sample size (ie, α=1), then the sample size may be doubled for each successive exploration of that configuration. Similarly, a stepwise test sample plan may be determined based on the functional dependence of the test time on the size of the sampled test data set.
[0054] beneficially, especially when combined with Figure 5 When used in combination with the planning method 500 and geometric sampling, the method 400 can significantly reduce the running time for determining the approximate optimal configuration. In some embodiments, as shown by experiments on multiple data sets, compared with an algorithm that determines the optimal configuration by training and testing each configuration on the entire training data set and the test data set, an approximate optimal configuration with an accuracy within 1% of the true optimal configuration can be obtained at a speed several tens of times faster. In addition, in the (unknown) function f i (l i ) and g i (u i ) can be assumed to be a convex function, when ∈ = 0 (which means that the running time of method 500 does not exceed four times the optimal running time), the four approximate guarantees of method 500 about the oracle optimal running time can be theoretically proved. As the loss tolerance ∈ increases, the running time tends to decrease significantly.
[0055] Of course, method 400 does not necessarily have to employ explicit planning method 500 and / or geometric sampling. Other methods for selecting configurations and associated sample sizes, as well as alternative methods for calculating confidence intervals, may be conceived by one of ordinary skill in the art, and these methods may retain some or all of the benefits of the specific embodiments described herein.
[0056] In general, the operations, algorithms, and methods described herein may be implemented in any suitable combination of software, hardware, and / or firmware, and the functionality provided may be grouped into multiple components, modules, or mechanisms. Modules and components may constitute software components (e.g., code embodied on a non-transitory machine-readable medium) or hardware-implemented components. Hardware-implemented components are tangible units that are capable of performing certain operations and may be configured or arranged in a certain manner. In example embodiments, one or more computer systems (e.g., stand-alone, client, or server computer systems) or one or more processors may be configured by software (e.g., an application or application portion) to operate as hardware-implemented components to perform certain operations described herein.
[0057] In various embodiments, hardware-implemented components may be implemented mechanically or electronically. For example, hardware-implemented components may include dedicated circuitry or logic that is permanently configured (e.g., as a dedicated processor, such as a field programmable gate array (FPGA) or an application-specific integrated circuit (ASIC)) to perform certain operations. Hardware-implemented components may also include programmable logic or circuitry (e.g., as contained in a general-purpose processor or other programmable processor) that is temporarily configured by software to perform certain operations. It should be understood that the decision as to whether to mechanically implement a hardware-implemented component in a dedicated and permanently configured circuitry or in a temporarily configured circuitry (e.g., configured by software) may be determined based on cost and time considerations.
[0058] Accordingly, the term "hardware-implemented component" should be understood to include a tangible entity that is physically constructed, permanently configured (e.g., hardwired), or temporarily or provisionally configured (e.g., programmed) to operate in a certain manner and / or perform certain operations described herein. Considering embodiments in which hardware-implemented components are temporarily configured (e.g., programmed), each hardware-implemented component need not be configured or instantiated at any time. For example, where the hardware-implemented component includes a general-purpose processor configured using software, the general-purpose processor can be configured into correspondingly different hardware-implemented components at different times. The software can accordingly configure the processor to, for example, constitute a specific hardware-implemented component at one time and constitute a different hardware-implemented component at another time.
[0059] The components of hardware implementation can provide information to the components of other hardware implementations and receive information from the components of other hardware implementations. Therefore, the components of the hardware implementation described can be considered to be communication coupled. In the case of multiple such hardware-implemented components at the same time, communication can be realized by signal transmission (for example, by connecting the appropriate circuits and buses of the components of the hardware implementation). In the embodiment where multiple hardware-implemented components are configured or instantiated at different times, the communication between the components of such hardware implementations can be realized, for example, by storing and obtaining information in a memory structure accessible to the components of multiple hardware implementations. For example, a component of a hardware implementation can perform an operation, and the output of the operation is stored in a memory device coupled with it. Then, another component of hardware implementation can access the memory device later to retrieve and process the stored output. The components of hardware implementation can also initiate communication with input or output devices, and can operate on resources (for example, information collection).
[0060] The various operations of the exemplary methods described herein may be performed at least in part by one or more processors that are temporarily configured (e.g., configured by software) or permanently configured to perform the relevant operations. Whether temporarily or permanently configured, such processors may constitute processor-implemented components that operate to perform one or more operations or functions. In some example embodiments, the components referred to herein may include processor-implemented components.
[0061] Similarly, the methods described herein may be implemented at least in part by a processor. For example, at least some of the operations of the method may be performed by a processor or one of the components implemented by the processor. The execution of certain operations may be distributed between one or more processors, not only residing in a single machine, but also deployed across multiple machines. In some example embodiments, one or more processors may be located in a single location (e.g., in an office environment or server farm), while in other embodiments, the processors may be distributed in multiple locations.
[0062] The one or more processors may also operate to support execution of related operations in a "cloud computing" environment, or as "software as a service" (SaaS). For example, at least some operations may be performed by a group of computers (as an example of a machine including a processor), which may be accessed via a network (e.g., the Internet) and one or more appropriate interfaces (e.g., an application program interface (API)).
[0063] Example embodiments may be implemented in digital electronic circuitry, computer hardware, firmware or software, or a combination thereof. Example embodiments may be implemented using a computer program product, such as a computer program tangibly embodied in an information carrier (e.g., in a machine-readable medium) for execution by a data processing apparatus (e.g., a programmable processor, a computer or computers) or for controlling its operation.
[0064] A computer program may be written in any form of descriptive language, including compiled or interpreted languages, and may be deployed in any form, including as a stand-alone program or as a component, subroutine, or other unit suitable for a computing environment. A computer program may be deployed to be executed on one computer or on multiple computers, either at one site or distributed across multiple sites and interconnected by a communication network.
[0065] In example embodiments, operations may be performed by one or more programmable processors executing a computer program to perform functions by operating on input data and generating output. Method operations may also be performed by special purpose logic circuitry (e.g., FPGA or ASIC), and the apparatus of example embodiments may be implemented as special purpose logic circuitry.
[0066] A computing system may include a client and a server. The client and the server are usually remote from each other and usually interact through a communication network. The relationship between the client and the server is generated by a computer program that runs on a respective computer and has a client-server relationship with each other. In an embodiment where a programmable computing system is deployed, it will be understood that both hardware and software architectures are worthy of consideration. Specifically, it will be understood that the choice of implementing certain functions in a combination of permanently configured hardware (e.g., ASIC), temporarily configured hardware (e.g., a combination of software and a programmable processor), or permanently and temporarily configured hardware may be a design choice. In various example embodiments, the hardware (e.g., machine) and software architecture that can be deployed are listed below.
[0067] Figure 6It is a block diagram of a machine in the form of an example of a computer system 600, in which instructions 624 can be executed to cause the machine to perform any one or more of the methods discussed herein. In an alternative embodiment, the machine operates as a standalone device, or can be connected (e.g., networked) to other machines. In a networked deployment, the machine can operate as a server or client computer in a server-client network environment, or operate as a peer machine in a peer-to-peer (or distributed) network environment. The machine can be a personal computer (PC), a tablet computer, a set-top box (STB), a personal digital assistant (PDA), a cellular phone, a network device, a network router, a switch or a bridge, or any machine capable of executing instructions (sequentially or otherwise), which specify the actions to be performed by the machine. In addition, although only a single machine is shown, the term "machine" should also be understood to include any machine collection, which executes a set (or multiple sets) of instructions individually or collectively to perform any one or more methods discussed herein.
[0068] The example computer system 600 includes a processor 602 (e.g., a central processing unit (CPU), a graphics processing unit (GPU), or both), a main memory 604, and a static memory 606, which communicate with each other via a bus 608. The computer system 600 may also include a video display 610 (e.g., a liquid crystal display (LCD) or a cathode ray tube (CRT)). The computer system 600 also includes an alphanumeric input device 612 (e.g., a keyboard or a touch-sensitive display screen), a user interface (UI) navigation (or cursor control) device 614 (e.g., a mouse), a disk drive unit 616, a signal generating device 618 (e.g., a speaker), and a network interface device 620.
[0069] The disk drive unit 616 includes a machine-readable medium 622 having stored thereon one or more sets of data structures and instructions 624 (e.g., software) embodied or utilized by any one or more of the methodologies or functionality described herein. During execution of the instructions 624 by the computer system 600, the instructions 624 may also reside, in whole or in part, within the main memory 604 and / or the processor 602, wherein the main memory 604 and the processor 602 also constitute machine-readable media.
[0070] Although the machine-readable medium 622 is shown as a single medium in the example embodiment, the term "machine-readable medium" may include a single medium or multiple media (e.g., a centralized or distributed database and / or associated caches and servers) that store one or more instructions 624 or data structures. The term "machine-readable medium" should also be considered to include any tangible medium that can store, encode, or carry instructions 624 for execution by a machine and cause the machine to perform any one or more methods of the present disclosure, or can store, encode, or carry data structures used by or associated with such instructions 624. Therefore, the term "machine-readable medium" should be considered to include, but is not limited to, solid-state memory, and optical and magnetic media. Specific examples of machine-readable media 622 include non-volatile memory, such as semiconductor memory devices, such as erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), and flash memory devices; magnetic disks, such as internal hard disks and removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks.
[0071] Instructions 624 may be transmitted or received over a communication network 626 using a transmission medium. Instructions 624 may be transmitted using a network interface device 620 and any of a variety of well-known transmission protocols (e.g., HTTP). Examples of communication networks include a local area network (LAN), a wide area network (WAN), the Internet, a mobile phone network, a plain old telephone (POTS) network, and a wireless data network (e.g., Wi-Fi and 4G / 5G networks). The term "transmission medium" should be taken to include any intangible medium capable of storing, encoding, or carrying instructions 624 for execution by a machine, and includes digital or analog communication signals or other intangible media to facilitate communication of such software.
[0072] The following numbered examples are provided as illustrative embodiments.
[0073] Example 1: One or more machine-readable media storing instructions executed by one or more hardware processors, the execution of which causes the one or more hardware processors to determine an approximately optimal machine learning configuration in a set of machine learning configurations by performing operations, the operations comprising: selecting a machine learning configuration in the set for training and determining an associated sample size; causing the selected machine learning configuration to be trained on a sampled training data set having the associated sample size; estimating a confidence interval of a quality metric for the trained machine learning configuration; and pruning the set based on a comparison between the estimated confidence interval of the trained machine learning configuration and estimated confidence intervals of other machine learning configurations in the set.
[0074] Example 2: One or more machine-readable media according to Example 1, wherein the selecting operation, the causing operation, the estimating operation, and the pruning operation are performed iteratively.
[0075] Example 3: One or more machine-readable media according to Example 2, wherein the approximately optimal machine learning configuration is the last machine learning configuration remaining in the set during iterative pruning.
[0076] Example 4: One or more machine-readable media according to Example 2 or Example 3, wherein for each of the machine learning configurations, the associated sample size is gradually increased during repeated training iterations.
[0077] Example 5: One or more machine-readable media of Example 4, wherein the associated sample size increases geometrically upon repeated training iterations.
[0078] Example 6: One or more machine-readable media according to any one of Examples 1-5, wherein the confidence interval is estimated based at least in part on: a training value of the quality metric determined for the trained machine learning configuration on the sampled training data set and a test value of the quality metric determined for the trained machine learning configuration based on a sampled test data set.
[0079] Example 7: One or more machine-readable media of Example 6, wherein an upper bound of the estimated confidence interval is greater than the training value and a lower bound of the confidence interval is less than the test value.
[0080] Example 8: One or more machine-readable media according to any of Examples 1-7, wherein the quality metric measures the accuracy of predictions made by the trained machine learning configuration.
[0081] Example 9: One or more machine-readable media according to any one of Examples 1-8, wherein pruning the set of machine learning configurations includes determining a highest lower bound among lower bounds of the confidence intervals for the machine learning configurations within the set, and removing from the set any machine learning configuration whose upper bound of the confidence interval exceeds the highest lower bound by no more than a specified loss tolerance.
[0082] Example 10: One or more machine-readable media according to any of Examples 1-10, wherein the selection of a machine learning configuration for training is based at least in part on a training cost associated with reducing the confidence interval of the machine learning configuration within the set.
[0083] Example 11: One or more machine-readable media according to any of Examples 1-10, wherein the approximately optimal machine learning configuration is one of the one or more machine learning configurations remaining in the pruned set when a time limit has been reached.
[0084] Example 12: A method comprising: iteratively pruning a set of machine learning configurations based on a training dataset and a test dataset by using one or more hardware processors to perform operations, the operations comprising, in each iteration of multiple iterations: sampling the training dataset and the test dataset according to a sampling plan associated with a machine learning configuration selected from the set; training the selected machine learning configuration based on the sampled training dataset, and determining a training accuracy associated with the trained selected machine learning configuration; evaluating the trained selected machine learning configuration based on the sampled test dataset to determine a test accuracy associated with the trained selected machine learning configuration; determining a confidence interval associated with the trained selected machine learning configuration based at least in part on the training accuracy and the test accuracy; pruning the set of machine learning configurations based on a comparison between the determined confidence interval and confidence intervals associated with other machine learning configurations in the set; and selecting one of the machine learning configurations remaining in the pruned set for the next iteration.
[0085] Example 13: A method according to Example 12, wherein pruning the set of machine learning configurations includes: pruning the confidence interval with the highest lower bound among the confidence intervals associated with the machine learning configurations in the set with any machine learning configuration whose overlap with the confidence interval with the highest lower bound is no greater than a specified loss tolerance and whose confidence interval has the highest lower bound.
[0086] Example 14: A method according to Example 12 or Example 13, wherein the sampling plan associated with the machine learning configuration at least increases the sample size of the sampling training data set when repeatedly training the same machine learning configuration.
[0087] Example 15: A method according to any one of Examples 12-14, wherein selecting the machine learning configuration includes: for at least some iterations, identifying the machine learning configuration whose associated confidence interval has the highest upper bound, and if the training cost gradient of the identified machine learning configuration is lower than the sum of the training cost gradients of other machine learning configurations, then selecting the identified machine learning configuration, otherwise selecting the machine learning configuration whose associated confidence interval has the second highest upper bound.
[0088] Example 16: A method according to any one of Examples 12-15, wherein the confidence interval is also determined based on the sample sizes of the sampled training data set and the sampled test data set.
[0089] Example 17: A method according to any of Examples 12-16, wherein the set of machine learning configurations is iteratively pruned until the set of machine learning configurations includes only one remaining machine learning configuration.
[0090] Example 18: A system comprising: one or more hardware processors configured to implement multiple processing components for determining an approximately optimal machine learning configuration in a set of machine learning configurations, the processing components comprising: a training and testing component configured to: after selecting one of the machine learning configurations from the set, train the selected machine learning configuration on a sampled training data set, and calculate training and testing quality metrics associated with the trained machine learning configuration; and a sampling and planning component configured to calculate confidence intervals for the machine learning configuration from the training and testing quality metrics, to iteratively prune the set of machine learning configurations based on the confidence intervals, select a machine learning configuration for training by the training and testing component, and determine an associated sample size of the sampled training data set for the selected machine learning configuration.
[0091] Example 19: The system of Example 18, wherein the processing component further comprises: a data sampler configured to sample the training data set based on a sample size determined by the sampling and planning component for the selected machine learning configuration.
[0092] Example 20: A system according to Example 18 or Example 19, wherein the sampling and planning component determines the sample size of each machine learning configuration in the machine learning configuration based on a predetermined stepwise sampling plan associated with the machine learning configuration.
[0093] Although the embodiments have been described with reference to specific example embodiments, it is clear that various modifications and changes may be made to these embodiments without departing from the broader scope of the present disclosure. Therefore, the description and the drawings should be considered illustrative rather than restrictive. The drawings forming a part thereof illustrate specific embodiments in which the present subject matter can be practiced by way of illustration and not limitation. The illustrated embodiments are described in sufficient detail to enable those skilled in the art to practice the teachings disclosed herein. Other embodiments may be used and derived therefrom so that structural and logical substitutions and changes may be made without departing from the scope of the present disclosure. Therefore, this description should not be understood as limiting, and the scope of the various embodiments is limited only by the appended claims and the full scope of equivalents to which these claims are entitled.
Claims
1. One or more non-transitory machine-readable media storing instructions executed by one or more hardware processors, wherein execution of the instructions causes the one or more hardware processors to determine an approximately optimal machine learning configuration from a set of machine learning configurations by performing operations, wherein the operations include: selecting a machine learning configuration for training among the set of machine learning configurations and determining an associated sample size for training; causing the training data set to be sampled according to the determined sample size to obtain a sampled training data set; causing the selected machine learning configuration to be trained on the sample training data set to optimize a training value of a quality metric; causing the trained machine learning configuration to be tested on at least one sample of a test data set to determine a test value of the quality metric; estimating, based at least in part on the training value and the test value of the quality metric, a confidence interval for a true value of the quality metric if the selected machine learning configuration were trained and tested on the complete training dataset and the test dataset; as well as The set of machine learning configurations is pruned based on a comparison between the estimated confidence intervals of the trained machine learning configurations and estimated confidence intervals of other machine learning configurations in the set.
2. One or more machine-readable media according to claim 1, wherein the operations of selecting, causing a training data set to be collected, causing the selected machine learning configuration to be trained, causing the trained machine learning configuration to be tested, estimating, and pruning are performed iteratively.
3. One or more machine-readable media according to claim 2, wherein the approximately optimal machine learning configuration is the machine learning configuration that finally remains in the set of machine learning configurations after iterative pruning.
4. One or more machine-readable media according to claim 2, wherein for each of the machine learning configurations, the associated sample size is gradually increased in repeated training iterations.
5. The one or more machine-readable media of claim 4, wherein the associated sample size grows geometrically over repeated training iterations.
6. The one or more machine-readable media of claim 1, wherein the test value of the quality metric is determined for the trained machine learning configuration based on a sample test data set.
7. The one or more machine-readable media of claim 1, wherein an upper bound of the estimated confidence interval is greater than the training value and a lower bound of the confidence interval is less than the test value.
8. The one or more machine-readable media of claim 1, wherein the quality metric measures the accuracy of predictions made by the trained machine learning configuration.
9. The one or more machine-readable media of claim 1, wherein pruning the set of machine learning configurations include: Determine a highest lower bound among the lower bounds of the confidence intervals for the machine learning configurations within the set of machine learning configurations, and remove from the set of machine learning configurations any machine learning configuration whose upper bound of the confidence interval exceeds the highest lower bound by no more than a specified loss tolerance.
10. One or more machine-readable media according to claim 1, wherein the selection of a machine learning configuration for training is based at least in part on a training cost associated with reducing the confidence interval of the machine learning configuration within the set of machine learning configurations.
11. One or more machine-readable media according to claim 1, wherein the approximately optimal machine learning configuration is one of the one or more machine learning configurations that remain within the pruned set when a time limit is reached.
12. A method, include: Iteratively pruning a set of machine learning configurations based on a training dataset and a test dataset by performing operations using one or more hardware processors, the operations comprising, in each iteration of a plurality of iterations: sampling the training dataset and the test dataset according to a sampling plan associated with a machine learning configuration selected from the set of machine learning configurations; training the selected machine learning configuration based on a sample training data set, and determining a training accuracy associated with the trained, selected machine learning configuration; evaluating the trained, selected, machine learning configuration based on a sample test data set to determine a test accuracy associated with the trained, selected, machine learning configuration; determining, based at least in part on the training accuracy and the testing accuracy, a confidence interval associated with the trained, selected machine learning configuration, the confidence interval providing an estimated bound on a true test accuracy if the selected machine learning configuration were trained and tested on the complete training and testing datasets; pruning the set of machine learning configurations based on a comparison between the determined confidence interval and confidence intervals associated with other machine learning configurations within the set of machine learning configurations; as well as One of the machine learning models among the remaining machine learning configurations in the pruned set is selected for the next iteration.
13. The method of claim 12, wherein pruning the set of machine learning configurations include: Among the confidence intervals associated with the machine learning configurations within the set of machine learning configurations, the confidence interval with the highest lower bound is compared to all other confidence intervals, and any machine learning configuration whose associated confidence interval overlaps the confidence interval with the highest lower bound by no more than a specified loss tolerance is removed from the set of machine learning configurations.
14. The method of claim 12, wherein the sampling plan associated with the machine learning configuration at least increases the sample size of the sampled training data set in repeated training of the same machine learning configuration.
15. The method of claim 12, wherein the machine learning configuration is selected include: For at least some iterations, a machine learning configuration having a highest upper bound of its associated confidence interval is identified, and if a training cost gradient of the identified machine learning configuration is lower than a sum of training cost gradients of the other machine learning configurations, the identified machine learning configuration is selected, otherwise a machine learning configuration having a second highest upper bound of its associated confidence interval is selected.
16. The method of claim 12, wherein the confidence interval is further determined based on sample sizes of the sampled training data set and the sampled test data set.
17. The method of claim 12, wherein the set of machine learning configurations is iteratively pruned until the set of machine learning configurations includes only one remaining machine learning configuration.
18. A system, include: One or more hardware processors configured to implement a plurality of processing components for determining an approximately optimal machine learning configuration from a set of machine learning configurations, the processing components comprising: A training and testing component configured to: after selecting one of the machine learning configurations from the set of machine learning configurations and determining a sample size associated with the machine learning configuration, causing a training data set to be sampled according to the sample size associated with the machine learning configuration to obtain a sampled training data set; causing the selected machine learning configuration to be trained on the sample training data set to optimize a training value of a quality metric; causing the trained machine learning configuration to be tested on at least one sample of a test data set to determine a test value of the quality metric; and The sampling and planning components are configured as follows: calculating, from the training value and the test value of the quality metric, a true value of the quality metric if the selected machine learning configuration were trained and tested on the complete training dataset and the test dataset; and The set of machine learning configurations is iteratively pruned based on the confidence interval, a machine learning configuration is selected for training by the training and testing components, and an associated sample size for the sampled training data set is determined for the selected machine learning configuration.
19. The system of claim 18, wherein the processing component further include: A data sampler is configured to sample the training data set based on the sample size associated with the selected machine learning configuration.
20. A system according to claim 18, wherein the sampling and planning component determines the sample size associated with each machine learning configuration based on a predetermined stepwise sampling plan associated with each of the machine learning configurations.
Citation Information
Patent Citations
Resource allocation for machine learning
US20140172753A1