Machine Learning Hyperparameter Tuning
The automated hyperparameter tuning system addresses inefficiencies in conventional methods by using a hyperparameter controller to efficiently determine optimal hyperparameters for machine learning models, reducing training time and resource use through parallel model training and performance-based selection.
Patent Information
- Application Number
- JP2023571357
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-05-17
- Filing Date
- 2022-05-15
- Publication Date
- 2025-06-25
- Estimated Expiration
- 2042-05-15
AI Technical Summary
Conventional machine learning hyperparameter tuning is a manual trial-and-error process that requires significant time and resources, as hyperparameters cannot be inferred during model fitting, leading to inefficiencies.
A method and system for automated hyperparameter tuning using a hyperparameter controller that performs operations on data processing hardware, including receiving optimization requests, determining hyperparameter permutations, training models, and selecting the best model based on performance, leveraging cloud computing and batch Gaussian process bandit optimization.
Automates hyperparameter tuning, reducing training time and resource consumption by training multiple models in parallel and providing efficient selection of optimal hyperparameters, thereby optimizing the machine learning process.
Smart Images

Figure 0007698740000001 
Figure 0007698740000002 
Figure 0007698740000003
Abstract
Description
Technical Field
[0001] Technical Field This disclosure relates to machine learning hyperparameter tuning.
Summary of the Invention
Problems to be Solved by the Invention
[0002] Background Machine learning hyperparameters are values used to control the learning process of a machine learning model. For example, machine learning hyperparameters include the topology of the model, the size of the model, and the learning rate of the model. Hyperparameter tuning has conventionally been a manual trial-and-error attempt because hyperparameters cannot be inferred while the model is being fitted to the training data. Thus, in conventional machine learning models, a significant portion of the time and resources may be required to perform sophisticated, manual, and laborious deliberations aimed at exploring or determining the optimal hyperparameters.
Means for Solving the Problems
[0003] Overview One aspect of the disclosure provides a method executed by a computer for performing machine learning hyperparameter tuning that causes the data processing hardware to perform operations when executed by the data processing hardware. The operations include receiving, from a user device, a hyperparameter optimization request requesting optimization of one or more hyperparameters of a machine learning model. The operations also include obtaining training data for training the machine learning model and determining a set of hyperparameter permutations of one or more hyperparameters of the machine learning model. For each respective hyperparameter permutation within the set of hyperparameter permutations, the operations include training a unique machine learning model using the training data and the respective hyperparameter permutation and determining the performance of the trained unique machine learning model. The operations also include selecting, based on the performance of each of the trained unique machine learning models, one of the trained unique machine learning models. The operations include generating one or more predictions using the selected one of the trained unique machine learning models.
[0004] Implementations of the disclosure may include one or more of the following optional features. In some implementations, the operations include determining a set of hyperparameter permutations that includes performing a search in a hyperparameter search space of one or more hyperparameters of the machine learning model. In some of these implementations, the operations include performing the search using batch Gaussian process bandit optimization. Optionally, the operations include determining a set of hyperparameter permutations based on one or more hyperparameters of the machine learning model and one or more previously trained machine learning models each sharing at least one hyperparameter. The one or more previously trained machine learning models may be associated with a user of the user device.
[0005] In some examples, training a unique machine learning model includes training two or more unique machine learning models in parallel. Optionally, providing the performance of each of the trained unique machine learning models to a user device includes providing an indication to the user device indicating which trained unique machine learning model has the best performance based on the training data. A hyperparameter optimization request may include an SQL query. Optionally, the hyperparameter optimization request includes a budget, and the size of the hyperparameter permutations of one or more hyperparameters of the machine learning model is based on the budget. In some examples, the data processing hardware is part of a distributed computing database system. In another implementation, selecting one of the trained unique machine learning models includes sending the performance of each of the trained unique machine learning models to the user device and receiving from the user device a trained unique machine learning model selection selecting one of the trained unique machine learning models.
[0006] Another aspect of the disclosure provides a system for performing machine learning hyperparameter tuning. The system includes data processing hardware and memory hardware that communicates with the data processing hardware. The memory hardware stores instructions that, when executed on the data processing hardware, cause the data processing hardware to perform operations. The operations include receiving, from a user device, a hyperparameter optimization request that requests optimization of one or more hyperparameters of a machine learning model. The operations also include obtaining training data for training the machine learning model and determining a set of hyperparameter permutations of one or more hyperparameters of the machine learning model. For each respective hyperparameter permutation within the set of hyperparameter permutations, the operations include training a unique machine learning model using the training data and the respective hyperparameter permutation and determining the performance of the trained unique machine learning model. The operations also include selecting, based on the performance of each of the trained unique machine learning models, one of the trained unique machine learning models. The operations include generating one or more predictions using the selected one of the trained unique machine learning models.
[0007] This aspect may include one or more of the following optional features. In some implementations, the operation includes determining a set of hyperparameter permutations that includes performing a search in a hyperparameter search space of one or more hyperparameters of a machine learning model. In some of these implementations, the operation includes performing the search using batch Gaussian process bandit optimization. Optionally, the operation includes determining a set of hyperparameter permutations based on one or more hyperparameters of a machine learning model and one or more previously trained machine learning models each sharing at least one hyperparameter. The one or more previously trained machine learning models may be associated with a user of a user device.
[0008] In some examples, training a unique machine learning model includes training two or more unique machine learning models in parallel. Optionally, providing the performance of each of the trained unique machine learning models to a user device includes providing an indication to the user device indicating which of the trained unique machine learning models has the best performance based on training data. The hyperparameter optimization request may include an SQL query. Optionally, the hyperparameter optimization request includes a budget, and the size of the hyperparameter permutations of one or more hyperparameters of the machine learning model is based on the budget. In some examples, the data processing hardware is part of a distributed computing database system. In another implementation, selecting one of the trained unique machine learning models includes sending the performance of each of the trained unique machine learning models to the user device and receiving from the user device a selection of one of the trained unique machine learning models.
[0009] Details of one or more implementations of the disclosure are set forth in the accompanying drawings and the description below. Other aspects, features, and advantages will be apparent from the description and drawings, and from the claims.
Brief Description of the Drawings
[0010]
Figure 1
Figure 2
Figure 3A
Figure 3B
Figure 4
Figure 5
Best Mode for Carrying Out the Invention
[0011] In the various drawings, like reference numerals represent like elements. Detailed Description Machine learning hyperparameters are values used to control the learning process of a machine learning model. For example, machine learning hyperparameters include the topology of the model, the size of the model, and the learning rate of the model. Hyperparameter tuning is a manual trial-and-error attempt in the conventional approach because hyperparameters cannot be inferred while the model is fitting to the training data. Thus, in a conventional machine learning model, a significant portion of time and resources may be required for determining and / or exploring the optimal hyperparameters. Thus, it is advantageous to incorporate a controller that can fully or partially automate hyperparameter tuning and the training of the machine learning model (i.e., reduce or eliminate manual tuning), and the efficiency can be further optimized by leveraging a cloud computing system.
[0012] Implementations herein include a hyperparameter controller that performs automatic hyperparameter tuning between distributed computing systems (e.g., cloud database systems). The controller can implement an interface based on a structured query language (SQL) that enables a user to automate hyperparameter tuning within a cloud computing system, and the search algorithm can automatically search for optimal hyperparameters for training a machine learning model. For example, the controller may include a search space for automatic hyperparameter search for use during training of a machine learning model.
[0013] In addition, the controller can collect and apply previously trained models to perform training of future models. This maximizes the efficiency of the system by utilizing previously stored information to update and train new models within the system. An automated process that explores and applies optimized hyperparameters maximizes efficiency for the user, freeing the user from having to manually explore and perform comparisons at the individual model level. The system can train multiple models in a single iteration (i.e., in parallel) to greatly reduce training time. The system can provide the user with the performance of each trained model, and in some instances, can automatically select the best model from each of the automatically trained models.
[0014] Referring now to FIG. 1, in some implementations, an exemplary hyperparameter tuning system 100 includes a remote system 140 that communicates with one or more user devices 10 via a network 112. The remote system 140 may be a distributed system (e.g., a cloud environment) having a scalable / elastic resource 142 that includes a single computer, multiple computers, or computing resources 144 (e.g., data processing hardware) and / or storage resources 146 (e.g., memory hardware). A data store 150 (i.e., a remote storage device) may be overlaid on the storage resource 146 to enable scalable use of the storage resource 146 by one or more of the clients (e.g., user devices 10) or the computing resources 144. The data store 150 is configured to store training data 152 (e.g., within a cloud database). The training data 152 may be associated with or controlled by the user 12.
[0015] The remote system 140 is configured to receive a hyperparameter optimization request 20 from a user device 10 associated with each respective user 12 via, for example, a network 112. The user device 10 may correspond to any computing device such as a desktop workstation, a laptop workstation, or a mobile device (i.e., a smartphone). The user device 10 includes computing resources 18 (e.g., data processing hardware) and / or storage resources 16 (e.g., memory hardware). The user 12 can create the request 20 using a Structured Query Language (SQL) interface 14. That is, the user 12 can generate a hyperparameter optimization request 20 using an SQL query. Each hyperparameter optimization request 20 requests the remote system 140 to optimize one or more hyperparameters 22, 22a - n of a machine learning model 210.
[0016] The remote system 140 causes the hyperparameter controller 160 to execute, which receives a request 20 that requests the hyperparameter controller 160 to optimize one or more hyperparameters 22 of the machine learning model 210 and to train the model 210 using the optimized hyperparameters 22. Each hyperparameter 22 for the hyperparameter controller 160 for tuning has a plurality of possible values that may be used to train the machine learning model 210. Some possible values of these hyperparameters 22 are more optimal than other possible hyperparameter values 22 (e.g., result in a faster or more efficient training process).
[0017] The hyperparameter controller 160 includes a permutation controller 230 that receives the request 20 and obtains the hyperparameters 22. The request can identify some or all of the hyperparameters 22 for tuning. Additionally or alternatively, the permutation controller 230 obtains one or more default hyperparameters 22 that are not identified by the request 20. The permutation controller 230 generates or determines a set of hyperparameter permutations 232, 232a - n based on the hyperparameters 22. Each hyperparameter permutation 232 includes different values for at least one of the hyperparameters 22. Using an example simplified for clarity, if the permutation controller 230 receives three hyperparameters 22 each having possible values of 1, 2, or 3, the permutation controller 230 can generate a first hyperparameter permutation 232 having the value {1,1,1}, a second hyperparameter permutation 232 having the value {1,1,2}, a third hyperparameter permutation 232 having the value {1,1,3}, a fourth hyperparameter permutation 232 having the value {1,2,1}, and so on. The set of hyperparameter permutations 232 includes some or all of the different combinations of possible values for the hyperparameters 22 of the machine learning model 210.
[0018] The permutation controller 230 can determine a set of hyperparameter permutations 232 using one or more tuning algorithms (i.e., can tune hyperparameters 22). One or more of the tuning algorithms may be default and / or selected by the user 12 (e.g., via the request 20). The tuning algorithms may be used to tune hyperparameters 22, 22a - n used to train the machine learning model 210 (i.e., to adjust the values). In some implementations, the permutation controller 230 determines whether the hyperparameter 22 is valid or invalid. When the permutation controller 230 determines that the hyperparameter 22 is invalid (e.g., an invalid value, incompatible with other hyperparameters 22 or the model 210, etc.), then the permutation controller 230 can discard or otherwise not use the hyperparameter permutation 232 that includes the invalid hyperparameter 22.
[0019] In some examples, the request 20 includes a plurality of hyperparameter permutations 232 to be generated (or, as discussed in more detail below, a plurality of machine learning models 210 to be trained). That is, the request can include a training budget. The permutation controller 230 can stop generating hyperparameter permutations 232 when the budget is reached. For example, the request 20 indicates that the user 12 desires that the maximum number of hyperparameter permutations 232 to be generated is 100.
[0020] The hyperparameter controller 160 also includes a model trainer 240. The model trainer 240 obtains training data 152 for training the machine learning model 210. The model trainer 240 can retrieve the training data 152 from, for example, the data store 150. In other examples, the request 20 includes the training data 152. The training data 152 can include any type of data on which the machine learning model 210 is trained to receive (e.g., text, images, audio, etc.). For example, the training data 152 includes data from a database, and the machine learning model 210 is trained to predict future values based on the values from the database. The model trainer 240 also receives a set of hyperparameter permutations 232 (i.e., different combinations of different values for each of the hyperparameters 22).
[0021] For each of the hyperparameter permutations 232 within the set of hyperparameter permutations 232, the model trainer 240 can train unique machine learning models 210, 210a - n using the training data 152 and each of the hyperparameter permutations 232. For example, when there are 50 different hyperparameter permutations 232, the model trainer 240 trains 50 different machine learning models 210 (i.e., one for each of the 50 different hyperparameter permutations 232). In some examples, claim 20 limits or restricts the number of models 210 to be trained to a number less than the total number of hyperparameter permutations 232. Each machine learning model 210 may be trained using the same training data 152 with the hyperparameters 22 required by the corresponding hyperparameter permutation 232. That is, each machine learning model 210 is trained using the same training data 152 but with different values for the hyperparameters 22. The model trainer 240 can train two or more of the machine learning models 210 in parallel (i.e., simultaneously) as described in more detail below. Alternatively, the model trainer 240 can train the models 210 serially.
[0022] Referring now to FIG. 2, the permutation controller 230 of the hyperparameter controller 160 determines a set of hyperparameter permutations 232 from the hyperparameter search space 234 (i.e., by exploring the hyperparameter search space 234). The hyperparameter search space 234 represents the available area that defines the set of all possible solutions for hyperparameter 22 tuning. For example, using 10 hyperparameters 22 each having 100 possible values, the hyperparameter search space 234 is a total of 100 10includes a number of possible solutions. As the number of hyperparameters 22 increases, it is readily apparent that the hyperparameter search space 234 grows rapidly to an unmeasurable size. Thus, the permutation controller 230 can attempt to "reduce" the hyperparameter search space 234 intelligently or efficiently by discarding known insufficient parts and / or focusing on known effective parts.
[0023] In some implementations, the permutation controller 230 determines a set of hyperparameter permutations 232 based at least in part on a model 210 previously trained by the model trainer 240. As shown by the schematic diagram 200, the permutation controller 230 determines one or more models 210, and the model trainer 240 selects or provides one or more of the models 210 that were previously trained for the user 12 (e.g., via the profile or identification information of the user 12) and / or for which the user 12 was previously trained (e.g., via the request 20). In these implementations, the previously trained model 210 is associated with the user 12 of the user device 10. In other examples, the permutation controller 230 selects a model 210 that was previously trained along with training data 152 similar to the current training data. Regardless of the source, the permutation controller 230 can determine the hyperparameter permutations 232 using the hyperparameters 22 selected for the previously trained model 210 as a guide. For example, the permutation controller 230 determines a set of hyperparameter permutations 232 based on one or more previously trained machine learning models 210 that each share at least one hyperparameter 22 with the hyperparameters 22 of the current machine learning model 210 and / or the request 20. The permutation controller 230 can use the hyperparameters 22 of the previously trained machine learning model 210 to reduce the hyperparameter search space 234 by freezing or restricting the values of the hyperparameters 22 that are along the hyperparameters of the previously trained machine learning model 210. The permutation controller 230 can retrieve hyperparameters from the data store 150.Similarly, once the training of the machine learning model 210 is completed, one or more hyperparameters 22 of the trained model may be stored in a hyperparameter table or other data structure at the data store 150. The table may be updated as the model trainer 240 trains new machine learning models 210.
[0024] In some implementations, the permutation controller 230 uses transfer learning to improve the selection of hyperparameters 22. In these implementations, the permutation controller 230 includes at least a subset of the same hyperparameters 22 to improve the search for optimal hyperparameters 22 and leverages data from previously trained (i.e., trained before receiving the current optimization request 20) machine learning models 210 associated with the user 12. Transfer learning may help avoid a "cold start" where an initial batch of hyperparameters 22 is selected via random exploration. As discussed above, the previously trained machine learning model 210 may be associated with the same user 12 who provided the current optimization request 20. In other examples, the previously trained machine learning model 210 is not associated with the same user 12.
[0025] In some implementations, the permutation controller 230 uses an algorithm to automatically find or search for optimal hyperparameters 22 within a hyperparameter search space 234 (e.g., based on Gaussian process bandit, covariance matrix adaptation evolution strategy, random search, grid search, etc.).
[0026] In some examples, user 12 provides restrictions to hyperparameter search space 234 via request 20. For example, request 20 may include restrictions on the values of one or more hyperparameters 22 or limitations on permutation controller 230 for a particular algorithm. When request 20 does not provide such limitations, permutation controller 230 can apply one or more default limitations to hyperparameter search space 234. Additionally or alternatively, permutation controller 230 supports conditional hyperparameters 22 that are only applicable when given that certain conditions are satisfied.
[0027] In some examples, permutation controller 230 begins hyperparameter tuning by solving a black-box optimization problem, i.e., finding X that optimizes the "black-box" objective function f: X -> R. By "black-box", it is only possible to observe the function output for a given finite input time with a relatively high cost, and other information about the function f, such as the gradient and Hessian of the function, cannot be utilized. In some implementations, although the controller uses Gaussian process bandit as the default algorithm to solve the above black-box optimization problem, other algorithms (e.g., covariance matrix adaptation evolution strategy, random search, grid search, etc.) may also be the default. Request 20 may override the default algorithm by specifying a particular algorithm and / or by providing an external algorithm. When the function f is modeled as a Gaussian process parameterized by x, or more specifically, as f(x) ~ GP(u(x), k(x,x')) with mean u(x) and covariance k(x,x'), the controller can solve it using Gaussian process regression fitting. *
[0028] In some examples, given past observations of pairs: (x_1,f(x_1)),(x_2,f(x_2)),...,(x_t,f(x_t)), the permutation controller 230 fits and / or updates a Gaussian process model (Gaussian process regressor) parameterized using the past observations. The permutation controller 230 can suggest x_t+1 using a Bayesian sampling procedure, and an exploration / exploitation balance strategy for a multi-armed bandit problem (i.e., for x) that maximizes both the mean and variance of the modeled f(x) will be selected as x_t+1 with the highest probability.
[0029] Referring now to FIGS. 3A and 3B, the user can set or specify or request the total number of models 210 (and the number of models 210 trained in parallel) trained based on, for example, budget 320 provided via claim 20. Budget 320 may correspond to the number of attempts requested for user 12 to perform, a monetary value related to the cost of operating or utilizing remote system 140, the number of models 210 that user 12 selects to be trained, and / or other aspects where user 12 may set parameters. For example, as depicted in FIG. 3A, user 12 sets an increase in budget 320 that results in permutation controller 230 generating five hyperparameter permutations 232, 232a - e. The number of hyperparameter permutations 232 directly corresponds to the number of models 210 that model trainer 240 trains in this example. Schematic diagram 300a includes model trainer 240 training five models 210, 210a - e corresponding to the five hyperparameter permutations 232, 232a - e received from permutation controller 230. Continuing the example of FIG. 3A, schematic diagram 300b (FIG. 3B) illustrates user 12 decreasing budget 320 so that fewer models 210 are trained. Here, the decreasing budget 320 results in two hyperparameter permutations 232a, 232b generated by permutation controller 230. As a result, model trainer 240 trains two models 210a, 210b. Budget 320 may be adjusted according to the user 12's computing parameters such that more than five models 210, 210a - n can be trained using more than five permutations 232, 232a - n and / or a single model 210 can be trained using a single hyperparameter permutation 232.These are simplified examples, and a remote system can generate hundreds, thousands, or even millions of different hyperparameter permutations 232.
[0030] The number of models 210 trained by the model trainer 240 may be directly related to the number of hyperparameter permutations 232 from the permutation controller 230. The budget 320 can thus dictate the number of models 210 to be trained by dictating the number of hyperparameter permutations 232 determined by the permutation controller 230. Put another way, the user 12 can adjust the number of models 210 generated by adjusting the size of the budget 320. Additionally or alternatively, the budget 320 may be used to determine the size (e.g., duration, amount of resources to expend, etc.) of the hyperparameter search space 234 to be explored such that a default amount of exploring the hyperparameter search space 234 is selected based on the budget 320. For example, the hyperparameter controller 160 tunes the hyperparameters 22 based on the priority of the models 210 within the allotted budget 320.
[0031] Referring back to FIG. 1, once the model trainer 240 trains the model 210, the performance controller 180 determines the respective performances 182, 182a - n of each trained model 210. For example, the performance controller 180 uses some or all of the training data 152 to measure the accuracy of each model 210 by comparing the labels or annotations of the training samples with the predictions generated by each model 210. The performance controller 180 provides the determined performance 182 to the user device 10. The hyperparameter controller 160 can send the performance 182 along with other attributes of the model 210 (e.g., the size of the model 210). The user 12 can select one or more of the models 210 trained based on the provided performance 182 and / or other attributes. In some examples, the hyperparameter controller 160 automatically selects a model 210 (e.g., the model 210 having the highest performance 182 or a model 210 that meets a default criterion or other pre - selected criterion). In these examples, the hyperparameter controller 160 can provide an indication to the user 12 for whom the model 210 is selected. In some implementations, in addition to the performance 182, the performance controller 180 provides an indication 184 that the trained model 210 has the highest performance 182 based on the training data 152 (i.e., by transmitting via the network 112). The user 12 can further determine which of the models 210 trained based on the indication 184 and any other attributes provided by the hyperparameter controller 160 to select.
[0032] User 12 can select one of the machine learning models 210 trained by sending the trained model selection 172 to the prediction generator 170 of the hyperparameter controller 160. In other examples, the performance controller 180 sends the trained model selection 172 to the prediction generator 170. The prediction generator 170 generates a prediction 174 based on the model selection 172 received from the user device 10. For example, the prediction generator 170 receives additional data (e.g., via the data store 150 or via the user device 10), and the selected model 210 makes one or more predictions based on the additional data. The prediction 174 may be provided to the user device 10. Alternatively, the hyperparameter controller 160 may bypass the user device and simply generate a trained model selection 172 that directly selects one of the trained unique models 210 having the best performance 182, and then directly provide the trained model selection 172 to the prediction generator 170 to generate the prediction 174.
[0033] FIG. 4 is a flowchart of an exemplary arrangement of operations for a method 400 of tuning hyperparameter 22. The method 400 executed by a computer causes the data processing hardware 144 to perform operations when executed by the data processing hardware 144. The method 400 includes, in operation 402, receiving, from the user device 10, a hyperparameter optimization request 20. The hyperparameter optimization request 20 requests optimization of one or more hyperparameters 22 of the machine learning model 210. The method 400 includes, in operation 404, obtaining training data 152 for training the machine learning model 210. The method 400 includes, in operation 406, determining a set of hyperparameter permutations 232 of the machine learning model 210. The method 400 includes, in operation 408, training a unique machine learning model 210 using the training data 152 and each hyperparameter permutation 232. The method 400 includes, in operation 410, determining the performance 182 of the trained unique machine learning model 210. The method 400 includes, in operation 412, selecting one of the trained unique machine learning models 210 based on the respective performance 182 of each of the trained unique machine learning models 210. In operation 414, the method 400 includes generating one or more predictions 174 using the selected one of the trained unique machine learning models 210.
[0034] FIG. 5 is a schematic diagram of an example computing device 500 that may be used to execute the systems and methods described in this document. Computing device 500 represents various forms of digital computers, such as a laptop, desktop, workstation, personal digital assistant, server, blade server, mainframe, and other appropriate computers. The components shown herein, their connections and relationships, and their functions are meant to be exemplary only and do not limit the implementations of the disclosure described and / or claimed in this document.
[0035] Computing device 500 includes a processor 510, a memory 520, a storage device 530, a high-speed interface / controller 540 connecting memory 520 and high-speed expansion port 550, and a low-speed interface / controller 560 connecting low-speed bus 570 and storage device 530. Each of the components 510, 520, 530, 540, 550, and 560 is interconnected using various buses and may be mounted on a common motherboard or in other manners as required. Processor 510 can process instructions for execution within computing device 500, including instructions stored in memory 520 or storage device 530 to display graphical information for a graphical user interface (GUI) on an external input / output device such as display 580 connected to high-speed interface 540. In other implementations, multiple processors and / or multiple buses may be used, along with multiple memories and multiple types of memory, as required. Also, multiple computing devices 500 may be connected (e.g., as a server bank, a group of blade servers, or a multi-processor system) such that each device provides a portion of the necessary operations.
[0036] Memory 520 stores information non-transitorily within computing device 500. Memory 520 may be a computer-readable medium, a volatile memory unit, or a non-volatile memory unit. The non-transitory memory 520 may be a physical device used to store programs (e.g., sequences of instructions) or data (e.g., program state information) either temporarily or permanently for use by computing device 500. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electrically erasable programmable read-only memory (EEPROM) (e.g., typically used for firmware such as a boot program). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase change memory (PCM), and disks or tapes.
[0037] Storage device 530 can provide mass storage for computing device 500. In some implementations, storage device 530 is a computer-readable medium. In various different implementations, storage device 530 may be a floppy disk device, a hard disk device, an optical disk device, or a tape device, a flash memory or other similar solid state memory device, or an array of devices including a storage area network or other configured devices. In further implementations, a computer program product is tangibly embodied in an information carrier. The computer program product includes instructions that, when executed, perform one or more methods such as those described above. The information carrier is a computer-readable medium or a machine-readable medium such as memory 520, storage device 530, or memory on processor 510.
[0038] The high-speed controller 540 manages the bandwidth-intensive operations for the computing device 500, while the low-speed controller 560 manages the lower-bandwidth-intensive operations. Such an assignment of duties is merely illustrative. In some implementations, the high-speed controller 540 is connected to the memory 520, to the display 580 (e.g., through a graphics processor or accelerator), and to the high-speed expansion port 550 into which various expansion cards (not shown) can be received. In some implementations, the low-speed controller 560 is connected to the storage device 530 and the low-speed expansion port 590. The low-speed expansion port 590, which can include various communication ports (e.g., USB, Bluetooth®, Ethernet®, wireless Ethernet), may be connected to one or more input / output devices such as a keyboard, a pointing device, a scanner, or to a networking device such as a switch or router, for example, via a network adapter.
[0039] As shown in the figure, the computing device 500 may be implemented in a plurality of different forms. For example, it may be implemented as a standard server 500a or multiple times in a group of such servers 500a, as a laptop computer 500b, or as part of a rack server system 500c.
[0040] The various implementations of the systems and techniques described herein can be realized in digital electronic circuits and / or optical circuits, integrated circuits, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations can be executed and / or interpreted on a programmable system including at least one programmable processor coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device, and can include implementation in one or more computer programs that are executable and / or interpretable by the programmable processor.
[0041] A software application (i.e., a software resource) may refer to computer software that causes a computing device to perform a task. In some examples, a software application may be referred to as an "application", an "app", or a "program". Example applications include, but are not limited to, system diagnostic applications, system management applications, system maintenance management applications, word processing applications, spreadsheet applications, messaging applications, media streaming applications, social networking applications, and gaming applications.
[0042] These computer programs (also known as programs, software, software applications or code) include machine instructions for a programmable processor and can be implemented in high-level procedural languages and / or object-oriented programming languages and / or assembly / machine languages. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any device and / or device (e.g., magnetic disk, optical disk, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including any computer program product, non-transitory computer-readable medium, machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0043] The processes and logical flows described in this specification can be performed by one or more programmable processors executing one or more computer programs to perform functions by processing input data and generating output, also referred to as data processing hardware. The processes and logical flows can also be performed by dedicated logic circuitry, such as, for example, an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit). Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, and any one or more processors of any kind of digital computer. In general, a processor will receive instructions and data from a read only memory or a random access memory or both. Essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. In general, a computer will also include, or be operatively coupled to receive data from, transfer data to, or both, one or more mass storage devices for storing data, such as, magnetic disks, magneto-optical disks, or optical disks. However, a computer need not necessarily have such devices. Computer readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including by way of example semiconductor memory devices, such as, EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD ROM disks and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, dedicated logic circuitry.
[0044] To provide interaction with a user, one or more aspects of the disclosure can be implemented on a computer having a display device for displaying information to the user, such as a CRT (cathode ray tube), LCD (liquid crystal display) monitor, or touch screen, and an optional keyboard and pointing device, such as a mouse or trackball, by which the user can provide input to the computer. Other types of devices can similarly be used to provide interaction with the user. For example, feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback, and input received from the user can be in any form, including acoustic input, speech input, or tactile input. Additionally, the computer can interact with the user by sending documents to and receiving documents from the devices used by the user, such as by sending a web page to a web browser on the user's client device in response to a request received from the web browser.
[0045] Numerous implementations have been described. Nevertheless, it will be understood that various modifications can be made without departing from the spirit and scope of the disclosure. Accordingly, other implementations are within the scope of the appended claims.
Claims
1. A method (400) executed by a computer that causes data processing hardware (144) to execute an operation when executed by the data processing hardware (144), the operation comprising: obtaining training data (152) for training a machine learning model (210); determining a set of hyperparameter permutations (232) of one or more hyperparameters (22) of the machine learning model (210); for each respective hyperparameter permutation (232) within the set of hyperparameter permutations (232), training a unique machine learning model (210) using the training data (152) and the respective hyperparameter permutation (232) based on the priority of the machine learning model (210); determining the performance (182) of the trained unique machine learning model (210); selecting one of the trained unique machine learning models (210) based on the respective performance (182) of each of the trained unique machine learning models (210).
2. The method (400) according to claim 1, wherein determining the set of hyperparameter permutations (232) comprises performing a search in a hyperparameter search space (234) of the one or more hyperparameters (22) of the machine learning model (210).
3. The method (400) according to claim 2, wherein performing the search in the hyperparameter search space (234) comprises performing the search using batch Gaussian process bandit optimization.
4. The method (400) according to any one of claims 1 to 3, wherein determining the set of hyperparameter permutations (232) is based on the one or more hyperparameters (22) of the machine learning model (210) and one or more previously trained machine learning models (210) each sharing at least one hyperparameter (22).
5. The machine learning model (210) trained in the one or more foregoing ways is the method (400) according to claim 4, associated with a user (12) of a user device (10). **Claim 6** Training the unique machine learning model (210) includes training two or more unique machine learning models (210) in parallel, the method (400) according to any one of claims 1 to 3. **Claim 7** Providing the performance of each of the trained unique machine learning models (210) to the user device (10) includes providing an indication to the user device (10) indicating which trained unique machine learning model (210) has the best performance (182) based on the training data (152), the method (400) according to any one of claims 1 to 3. **Claim 8** Further comprising receiving, from a user device (10), a hyperparameter optimization request (20) requesting optimization of the one or more hyperparameters (22) of the machine learning model (210), wherein the hyperparameter optimization request (20) includes an SQL query, the method (400) according to any one of claims 1 to 3. **Claim 9** A method (400) executed by a computer that causes a data processing hardware (144) to perform an operation when executed by the data processing hardware (144), the operation comprising: receiving, from a user device (10), a hyperparameter optimization request (20) requesting optimization of one or more hyperparameters (22) of a machine learning model (210); obtaining training data (152) for training the machine learning model (210); determining a set of hyperparameter permutations (232) of the one or more hyperparameters (22) of the machine learning model (210); for each respective hyperparameter permutation (232) within the set of hyperparameter permutations (232), training a unique machine learning model (210) using the training data (152) and the respective hyperparameter permutation (232); Determining the performance (182) of the trained unique machine learning model (210); selecting one of the trained unique machine learning models (210) based on the respective performance (182) of the trained unique machine learning models (210), wherein the hyperparameter optimization requirement (20) includes a budget (320), and the budget includes at least one of the number of optimization attempts, a monetary value related to the cost of utilizing the data processing hardware (144), the number of machine learning models (210), and the priority of the machine learning models (210); A method (400) wherein the size of the set of hyperparameter permutations (232) of the one or more hyperparameters (22) of the machine learning model (210) is based on the budget (320).
10. The method (400) according to any one of claims 1 to 3 and claim 9, wherein the data processing hardware (144) is part of a distributed computing database system (140).
11. Selecting the one of the trained unique machine learning models (210) comprises: transmitting the respective performance (182) of each of the trained unique machine learning models (210) to a user device (10); receiving, from the user device (10), a selection of the trained unique machine learning model (210) that selects one of the trained unique machine learning models (210), the method (400) according to any one of claims 1 to 3 and claim 9.
12. The method (400) according to any one of claims 1 to 3 and claim 9, further comprising generating one or more predictions (174) using the selected one of the trained unique machine learning models (210).
13. data processing hardware (144); A memory hardware (146) that communicates with the data processing hardware (144), the memory hardware (146) stores instructions, and when the instructions are executed on the data processing hardware (144), the instructions cause the data processing hardware (144) to execute operations, and the operations are obtaining training data (152) for training a machine learning model (210); determining a set of hyperparameter permutations (232) of one or more hyperparameters (22) of the machine learning model (210); for each of the hyperparameter permutations (22) within the set of hyperparameter permutations (232), training a unique machine learning model (210) using the training data (152) and the respective hyperparameter permutations (232) based on the priority of the machine learning model (210); determining the performance (182) of the trained unique machine learning model (210); A system (100) comprising selecting one of the trained unique machine learning models (210) based on the performance (182) of each of the trained unique machine learning models (210).
14. The determining the set of hyperparameter permutations (232) includes performing a search in a hyperparameter search space (234) of the one or more hyperparameters (22) of the machine learning model (210), according to the system (100) of claim 13.
15. Performing the search in the hyperparameter search space (234) includes performing the search using batch Gaussian process bandit optimization, according to the system (100) of claim 14.
16. The determining the set of hyperparameter permutations (232) is based on the one or more hyperparameters (22) of the machine learning model (210) and one or more previously trained machine learning models (210) each sharing at least one hyperparameter (22), according to the system (100) of any one of claims 13 to 15.
17. The system (100) according to claim 16, wherein the one or more pre-trained machine learning models (210) are associated with a user (12) of a user device (10).
18. The system (100) according to any one of claims 13 to 15, wherein training the unique machine learning model (210) includes training two or more of the unique machine learning models (210) in parallel.
19. The system (100) according to any one of claims 13 to 15, wherein providing the performance (182) of each of the trained unique machine learning models (210) to the user device (10) includes providing an indication to the user device (10) indicating which trained unique machine learning model (210) has the best performance (182) based on the training data (152).
20. The operation further includes receiving, from the user device (10), a hyperparameter optimization request (20) requesting optimization of the one or more hyperparameters (22) of the machine learning model (210), The system (100) according to any one of claims 13 to 15, wherein the hyperparameter optimization request (20) includes an SQL query.
21. Data processing hardware (144), Memory hardware (146) communicating with the data processing hardware (144), the memory hardware (146) storing instructions that, when executed on the data processing hardware (144), cause the data processing hardware (144) to perform operations, the operations including receiving, from the user device (10), a hyperparameter optimization request (20) requesting optimization of one or more hyperparameters (22) of a machine learning model (210); obtaining training data (152) for training the machine learning model (210); determining a set of hyperparameter permutations (232) of the one or more hyperparameters (22) of the machine learning model (210); For each of the hyperparameter permutations (22) within the set of the hyperparameter permutations (232), training a unique machine learning model (210) using the training data (152) and the respective hyperparameter permutation (232); determining the performance (182) of the trained unique machine learning model (210); selecting one of the trained unique machine learning models (210) based on the performance (182) of each of the trained unique machine learning models (210), wherein the hyperparameter optimization request (20) includes a budget (320), and the budget includes at least one of a monetary value related to the cost of utilizing the data processing hardware (144), the number of the machine learning models (210), and the priority of the machine learning models (210), and the size of the set of the hyperparameter permutations (232) of the one or more hyperparameters (22) of the machine learning model (210) is based on the budget (320), for the system (100). **Claim 22** The system (100) according to any one of claims 13 to 15 and claim 21, wherein the data processing hardware (144) is part of a distributed computing database system (140). **Claim 23** The selecting of the one of the trained unique machine learning models (210) comprises transmitting the performance (182) of each of the trained unique machine learning models (210) to the user device (10); receiving, from the user device (10), a selection of the trained unique machine learning model (210) that selects the one of the trained unique machine learning models (210), for the system (100) according to any one of claims 13 to 15 and claim 21.
24. The system (100) according to any one of claims 13 to 15 and claim 21, wherein the operation further includes generating one or more predictions (174) using the selected one of the trained unique machine learning models (210).
Citation Information
Patent Citations
Information providing device and machine-readable recording medium where program is recorded
JP1999066101A
Search method, search device and search program
JP2019079214A
Hyper parameter tuning method, device and program
JP2020198135A
Distributed hyperparameter tuning system for machine learning
US20180240041A1
Fast hyperparameter search for machine-learning program
US20190122141A1