Hyperparameter optimization method, device, computing equipment, storage medium and program product
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-30
- Publication Date
- 2026-08-11
AI Technical Summary
[0003]现有技术中,通常是基于人工经验来配置机器学习模型的超参数,然而基于人工经验调优得到的超参数往往是次优解,并非最优超参数,并且基于此进行的超参数调优过程耗时长,适用性差
[0023]In the hyperparameter optimization method and apparatus according to some embodiments of this disclosure, an initialization step is first performed, followed by a confidence determination step, an optimal candidate step, and a judgment step according to predetermined logic. The initialization step includes obtaining a candidate hyperparameter set, which includes multiple candidate hyperparameters and performance metrics of a first portion of the candidate hyperparameters. The performance metrics indicate the performance of a machine learning model configured with the corresponding hyperparameters. The confidence determination step includes determining the confidence level that each hyperparameter outside the first portion of the candidate hyperparameters is the optimal hyperparameter based on the performance metrics of the first portion of the candidate hyperparameters. The optimal hyperparameter refers to the hyperparameter with the best performance metric. The optimal candidate step includes determining the optimal candidate hyperparameter among the non-first portion of the candidate hyperparameters based on the confidence level that each hyperparameter outside the first portion is the optimal hyperparameter, and obtaining the performance metrics of the optimal candidate hyperparameter. After executing the optimal candidate step, the process proceeds to the judgment step. The judgment step includes determining whether the performance index of the optimal candidate hyperparameter meets predetermined conditions. If the performance index does not meet the predetermined conditions, the optimal candidate hyperparameter is added to the first part of the hyperparameters. The performance index of the optimal candidate hyperparameter is added to the candidate hyperparameter set, and the process proceeds to the confidence determination step. If the performance index meets the predetermined conditions, the optimal candidate hyperparameter is determined as the optimal hyperparameter. Therefore, this disclosure avoids configuring hyperparameters of the machine learning model based on human experience by executing the above steps according to predetermined logic, achieving fast, efficient, and accurate hyperparameter optimization.
Smart Images

Figure CN117035003B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computers, and in particular to hyperparameter optimization methods, apparatus, computing devices, storage media, and program products. Background Technology
[0002] With the development of computer technology, machine learning models have been widely applied. Hyperparameters of a machine learning model refer to the parameters set before the learning process begins, rather than the parameters learned during training. Compared to parameters, hyperparameters define higher-level concepts of a machine learning model, such as model complexity or learning capacity. Therefore, selecting an optimal set of hyperparameters for a machine learning model is crucial for improving its learning performance and effectiveness.
[0003] In existing technologies, hyperparameters of machine learning models are usually configured based on human experience. However, hyperparameters tuned based on human experience are often suboptimal solutions, not optimal ones. Furthermore, the hyperparameter tuning process based on this is time-consuming and has poor applicability. Summary of the Invention
[0004] The inventors have noted that achieving fast, efficient, and accurate hyperparameter optimization is a pressing problem. Therefore, this application provides a hyperparameter optimization method, apparatus, computer-readable storage medium, and computer program product, aiming to solve the aforementioned problem.
[0005] According to a first aspect of this disclosure, a hyperparameter optimization method for a machine learning model is disclosed, characterized in that the method includes an initialization step, and a confidence determination step, an optimal candidate step, and a judgment step performed according to predetermined logic: Initialization step: Obtaining a candidate hyperparameter set, the candidate hyperparameter set including multiple candidate hyperparameters and performance indicators of a first portion of the candidate hyperparameters, the performance indicators being used to indicate the performance of the machine learning model configured with the corresponding hyperparameters; Confidence determination step: Based on the performance indicators of the first portion of the candidate hyperparameters, determining the confidence level that each hyperparameter is the optimal hyperparameter among the non-first portion of the candidate hyperparameters; The optimal hyperparameter refers to the hyperparameter with the best performance index. The optimal candidate step is as follows: Based on the confidence that each hyperparameter in the non-first part of the hyperparameters is the optimal hyperparameter, the optimal candidate hyperparameter in the non-first part of the hyperparameters is determined, and the performance index of the optimal candidate hyperparameter is obtained. The judgment step is as follows: Determine whether the performance index of the optimal candidate hyperparameter meets the predetermined conditions. If the performance index of the optimal candidate hyperparameter does not meet the predetermined conditions, the optimal candidate hyperparameter is added to the first part of the hyperparameters, and the performance index of the optimal candidate hyperparameter is added to the candidate hyperparameter set. Proceed to the confidence determination step. If the performance index of the optimal candidate hyperparameter meets the predetermined conditions, the optimal candidate hyperparameter is determined as the optimal hyperparameter.
[0006] According to a second aspect of this disclosure, a hyperparameter optimization apparatus for a machine learning model is disclosed. The hyperparameter optimization apparatus is configured to perform an initialization step, and execute a confidence determination step, an optimal candidate step, and a judgment step according to predetermined logic. The hyperparameter optimization apparatus includes: an initialization module configured to perform the initialization step: obtaining a candidate hyperparameter set, the candidate hyperparameter set including multiple candidate hyperparameters and performance indicators of a first portion of the candidate hyperparameters, the performance indicators being used to indicate the performance of a machine learning model configured with the corresponding hyperparameters; and a confidence module configured to perform the confidence determination step: based on the performance indicators of the first portion of the candidate hyperparameters, determining the confidence that each hyperparameter is the optimal hyperparameter for each hyperparameter not in the first portion of the candidate hyperparameters. The optimal hyperparameter refers to the hyperparameter with the best performance index. The optimal candidate module is configured to perform the optimal candidate step: determine the optimal candidate hyperparameter among the non-first part hyperparameters based on the confidence that each hyperparameter in the non-first part hyperparameters is the optimal hyperparameter, and obtain the performance index of the optimal candidate hyperparameter; and the judgment module is configured to perform the judgment step: determine whether the performance index of the optimal candidate hyperparameter meets a predetermined condition; in response to the performance index of the optimal candidate hyperparameter not meeting the predetermined condition, add the optimal candidate hyperparameter to the first part hyperparameters, and add the performance index of the optimal candidate hyperparameter to the candidate hyperparameter set, proceed to the confidence determination step; in response to the performance index of the optimal candidate hyperparameter meeting the predetermined condition, determine the optimal candidate hyperparameter as the optimal hyperparameter.
[0007] In a hyperparameter optimization apparatus according to some embodiments of this application, the confidence determination step further includes: performing a first transformation on the hyperparameters in the hyperparameter set to determine the hyperparameter transformation values of the hyperparameters, wherein the hyperparameter transformation values satisfy a normal distribution; performing a second transformation on the performance index of a first portion of the hyperparameters in the hyperparameter set to determine the performance index transformation values of the first portion of the hyperparameters, wherein the performance index transformation values satisfy a normal distribution; based on the hyperparameter transformation values of the first portion of the hyperparameters and their performance index transformation values, determining a first confidence level for each hyperparameter in the non-first portion of the hyperparameters that has the largest performance index transformation value; and determining the first confidence level of the hyperparameter transformation value of each hyperparameter as the confidence level that the corresponding hyperparameter is the optimal hyperparameter.
[0008] In a hyperparameter optimization apparatus according to some embodiments of this application, performing a first transformation on hyperparameters in a hyperparameter set and determining the hyperparameter transformation value includes: obtaining a normalized value of the hyperparameter, the normalized value being between 0 and 1; performing a first mapping on the normalized value of the hyperparameter to determine a first mapping value of the normalized value of the current hyperparameter; and determining the first mapping value of the normalized value of the hyperparameter as the hyperparameter transformation value of the current hyperparameter, wherein the normalized value of the hyperparameter and the first mapping value of the normalized value of the current hyperparameter satisfy the equation: Among them, h i It is the normalized value of the hyperparameter, H. i is the first mapping value of the normalized hyperparameters, a is the normalized value of the smallest hyperparameter in the hyperparameter set, and b is the normalized value of the largest hyperparameter in the hyperparameter set.
[0009] In a hyperparameter optimization apparatus according to some embodiments of this application, performing a second transformation on the performance index of a first portion of the hyperparameters in a hyperparameter set to determine the transformed performance index value of the first portion of the hyperparameters includes: obtaining the performance index of the hyperparameters from the hyperparameter set; performing a second mapping on the performance index of the hyperparameters to determine the second mapping value of the performance index of the hyperparameters; and determining the second mapping value of the performance index of the hyperparameters as the transformed performance index value of the hyperparameters, wherein the performance index of the hyperparameters and the second mapping value of the performance index of the hyperparameters satisfy the equation: Among them, P i It is the second mapping value of the performance index of the hyperparameter, p i These are performance metrics for hyperparameters. It is the mean of the performance metrics in the set of hyperparameters. It is the standard deviation of the performance metrics in the hyperparameter set.
[0010] In a hyperparameter optimization apparatus according to some embodiments of this application, determining the first confidence level of having the maximum performance index transformation value for each hyperparameter outside the first part of hyperparameters, based on the hyperparameter transformation values and performance index transformation values of the first part of hyperparameters, includes: establishing a statistical model based on the hyperparameter transformation values and performance index transformation values of the first part of hyperparameters, wherein the statistical model is used to determine the expectation and variance of the performance index transformation values of the hyperparameter transformation values; determining the expectation and variance of the performance index transformation values of each hyperparameter outside the first part of hyperparameters based on the statistical model; and determining the first confidence level of having the maximum performance index transformation value for the corresponding hyperparameter based on the expectation and variance of the performance index transformation values of the hyperparameter transformation values of each hyperparameter outside the first part of hyperparameters.
[0011] In a hyperparameter optimization apparatus according to some embodiments of this application, determining a first confidence level that the hyperparameter transformation value of a corresponding hyperparameter has the maximum performance index transformation value based on the expectation and variance of the performance index transformation value of the hyperparameter transformation value of each hyperparameter in the non-first part of hyperparameters includes: obtaining a confidence function, the confidence function being used to determine the confidence level based on the expectation and variance; for each hyperparameter transformation value in the non-first part of hyperparameters, substituting the expectation and variance of its performance index transformation value into the confidence function; and determining the output of the confidence function as the first confidence level that the hyperparameter transformation value of the corresponding hyperparameter has the maximum performance index transformation value.
[0012] In a hyperparameter optimization apparatus according to some embodiments of this application, obtaining the confidence function includes: obtaining a first acquisition function, a second acquisition function, and a third acquisition function; and determining a confidence function based on the first acquisition function, the second acquisition function, and the third acquisition function, wherein the confidence function, the first acquisition function, the second acquisition function, and the third acquisition function satisfy the following equation: in, It is a confidence function. It is the first acquisition function. It is the second acquisition function. It is the third acquisition function. It is the weight of the first acquisition function. It is the weight of the second acquisition function. Is it the weight of the third acquisition function and .
[0013] In the hyperparameter optimization apparatus according to some embodiments of this application, the first acquisition function, the second acquisition function, and the third acquisition function each include one or more of the following: expected increment acquisition function, probability increment acquisition function, and confidence upper bound acquisition function, and the weights of the first acquisition function, the second acquisition function, and the third acquisition function are all between 0 and 1 and are randomly generated.
[0014] In the hyperparameter optimization apparatus according to some embodiments of this application, the statistical model includes one or both of a Gaussian mixture model and a Gaussian model.
[0015] In a hyperparameter optimization apparatus according to some embodiments of the present application, obtaining a candidate hyperparameter set includes: obtaining a plurality of candidate hyperparameters; determining a first portion of hyperparameters among the plurality of candidate hyperparameters; obtaining performance indicators of the first portion of hyperparameters among the plurality of candidate hyperparameters; and establishing a candidate hyperparameter set, the candidate hyperparameter set including the plurality of candidate hyperparameters and the performance indicators of the first portion of hyperparameters among the plurality of candidate hyperparameters.
[0016] In a hyperparameter optimization apparatus according to some embodiments of the present application, determining a first portion of hyperparameters among a plurality of candidate hyperparameters includes: dividing the candidate hyperparameters into a first number of candidate hyperparameter groups, wherein each candidate hyperparameter group in the first number of candidate hyperparameter groups has an equal probability of containing the optimal hyperparameter; and randomly selecting a second number of candidate hyperparameters from each candidate hyperparameter group in the first number of candidate hyperparameter groups as the first portion of hyperparameters.
[0017] In a hyperparameter optimization apparatus according to some embodiments of this application, obtaining the performance index of a first portion of hyperparameters among a plurality of candidate hyperparameters includes: obtaining a training set for training a machine learning model; configuring a machine learning model with each of the first portion of hyperparameters; training the machine learning model with the training set to determine the parameters of the machine learning model; testing the performance of the trained machine learning model to determine the performance index of the trained machine learning model; and determining the performance index of the trained machine learning model as the performance index of the corresponding hyperparameter in the first portion of hyperparameters.
[0018] In some embodiments of the hyperparameter optimization apparatus according to this application, the predetermined conditions include: the performance index of the optimal candidate hyperparameter exceeds a predetermined threshold or the number of times the judgment step is executed exceeds a predetermined number.
[0019] In a hyperparameter optimization apparatus according to some embodiments of this application, obtaining the performance metric of the optimal candidate hyperparameter includes: obtaining a training set for training a machine learning model; configuring a machine learning model with each of the optimal candidate hyperparameters; training the machine learning model with the training set to determine the parameters of the machine learning model; testing the performance of the trained machine learning model to determine the performance metric of the trained machine learning model; and determining the performance metric of the trained machine learning model as the performance metric of the corresponding hyperparameter among the optimal candidate hyperparameters.
[0020] According to a third aspect of this disclosure, a computing device is disclosed, comprising: a memory configured to store computer-executable instructions; and a processor configured to perform any of the methods described above when the computer-executable instructions are executed by the processor.
[0021] According to a fourth aspect of this disclosure, a computer-readable storage medium is disclosed that stores computer-executable instructions, which, when executed, perform any of the methods described above.
[0022] According to a fifth aspect of this disclosure, a computer program product is disclosed, comprising computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, perform any of the methods described above.
[0023] In the hyperparameter optimization method and apparatus according to some embodiments of this disclosure, an initialization step is first performed, followed by a confidence determination step, an optimal candidate step, and a judgment step according to predetermined logic. The initialization step includes obtaining a candidate hyperparameter set, which includes multiple candidate hyperparameters and performance metrics of a first portion of the candidate hyperparameters. The performance metrics indicate the performance of a machine learning model configured with the corresponding hyperparameters. The confidence determination step includes determining the confidence level that each hyperparameter outside the first portion of the candidate hyperparameters is the optimal hyperparameter based on the performance metrics of the first portion of the candidate hyperparameters. The optimal hyperparameter refers to the hyperparameter with the best performance metric. The optimal candidate step includes determining the optimal candidate hyperparameter among the non-first portion of the candidate hyperparameters based on the confidence level that each hyperparameter outside the first portion is the optimal hyperparameter, and obtaining the performance metrics of the optimal candidate hyperparameter. After executing the optimal candidate step, the process proceeds to the judgment step. The judgment step includes determining whether the performance index of the optimal candidate hyperparameter meets predetermined conditions. If the performance index does not meet the predetermined conditions, the optimal candidate hyperparameter is added to the first part of the hyperparameters. The performance index of the optimal candidate hyperparameter is added to the candidate hyperparameter set, and the process proceeds to the confidence determination step. If the performance index meets the predetermined conditions, the optimal candidate hyperparameter is determined as the optimal hyperparameter. Therefore, this disclosure avoids configuring hyperparameters of the machine learning model based on human experience by executing the above steps according to predetermined logic, achieving fast, efficient, and accurate hyperparameter optimization.
[0024] These and other advantages of this disclosure will become clear from the embodiments described below, and will be illustrated with reference to the embodiments described below. Attached Figure Description
[0025] Embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings, in which: Figure 1 Exemplary application scenarios of hyperparameter optimization methods according to some embodiments of this application are illustrated; Figure 2 An exemplary flowchart of a hyperparameter optimization method according to some embodiments of this application is shown; Figure 3 An exemplary flowchart of a confidence level determination method according to some embodiments of this application is shown; Figure 4 An exemplary flowchart of a hyperparameter optimization method according to some embodiments of this application is shown; Figure 5 An exemplary schematic diagram of a hyperparameter optimization system according to some embodiments of this application is shown; Figure 6The performance of hyperparameter optimization methods according to some embodiments of this application is illustrated; Figure 7 The performance of hyperparameter optimization methods according to some embodiments of this application is illustrated; Figure 8 The performance of hyperparameter optimization methods according to some embodiments of this application is illustrated; Figure 9 An exemplary structural block diagram of a hyperparameter optimization apparatus according to some embodiments of this application is shown; and, Figure 10 An example system is shown, which includes an example computing device representing one or more systems and / or devices that can implement the various methods described herein. Detailed Implementation
[0026] The following description provides specific details of various embodiments of this disclosure to enable those skilled in the art to fully understand and implement the various embodiments of this disclosure. It should be understood that the technical solutions of this disclosure can be implemented without some of these details. In some cases, this disclosure does not show or describe in detail some well-known structures or functions to avoid such unnecessary descriptions that would obscure the description of the embodiments of this disclosure. The terminology used in this disclosure should be understood in its broadest and most reasonable manner, even when used in conjunction with specific embodiments of this disclosure.
[0027] First, some terms or expressions used in the embodiments of this application will be explained to facilitate understanding by those skilled in the art.
[0028] Hyperparameters: In the context of machine learning, hyperparameters are parameters whose values are set before the learning process begins, rather than parameters obtained through training. Typically, hyperparameters need to be optimized to select an optimal set for the learning machine, thereby improving learning performance and effectiveness.
[0029] Machine Learning (ML) is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and instructional learning.
[0030] Latin hypercube sampling (LHS) is a method of approximately random sampling from a multivariate parameter distribution. It belongs to stratified sampling techniques and is often used in computer experiments or Monte Carlo integration.
[0031] Gaussian model: A Gaussian model is a model that uses the Gaussian probability density function (normal distribution curve) to accurately quantify things, decomposing a thing into several models based on the Gaussian probability density function (normal distribution curve).
[0032] Gaussian Mixture Model: A Gaussian model uses the Gaussian probability density function (normal distribution curve) to accurately quantify things, decomposing a thing into several models based on the Gaussian probability density function (normal distribution curve).
[0033] Acquisition function: A function used in statistical learning to determine the next batch of sampling points, such as the expectation improvement function, the probability of improvement function, and the upper confidence bound function.
[0034] Figure 1 An exemplary application scenario 100 of the hyperparameter optimization method according to some embodiments of this application is shown. For example... Figure 1 As shown, application scenario 100 includes server 110, terminal 120, server 130, and network 140. Server 110, terminal 120, and server 130 are coupled together via network 140. As an example, hyperparameter optimization of a machine learning model can be performed on server 110, and then the optimal hyperparameters can be sent to terminal 120 or server 130 via server 140.
[0035] As an example, an initialization step can be performed on server 110, followed by a confidence determination step, an optimal candidate step, and a judgment step according to predetermined logic. The initialization step includes obtaining a candidate hyperparameter set, which includes multiple candidate hyperparameters and performance metrics for a first portion of the candidate hyperparameters. These performance metrics indicate the performance of the machine learning model configured with the corresponding hyperparameters. The confidence determination step includes determining the confidence level that each hyperparameter outside the first portion of the candidate hyperparameters is the optimal hyperparameter based on its performance metrics. The optimal candidate step includes determining the optimal candidate hyperparameter from the non-first portion of the candidate hyperparameters based on the confidence level that each hyperparameter outside the first portion is the optimal hyperparameter, and obtaining the performance metrics of the optimal candidate hyperparameter. The judgment step includes determining whether the performance index of the optimal candidate hyperparameter meets the predetermined conditions; in response to the performance index of the optimal candidate hyperparameter not meeting the predetermined conditions, the optimal candidate hyperparameter is added to the first part of hyperparameters; the performance index of the optimal candidate hyperparameter is added to the candidate hyperparameter set; and the process proceeds to the confidence determination step; in response to the performance index of the optimal candidate hyperparameter meeting the predetermined conditions, the optimal candidate hyperparameter is determined as the optimal hyperparameter.
[0036] As an example, the initialization steps can also be performed on server 130 or terminal 120, and then one or more of the confidence determination step, the optimal candidate step, and the judgment step can be performed according to predetermined logic.
[0037] In some embodiments, one or more of the above-described initialization steps, confidence determination steps, optimal candidate steps, and judgment steps may be performed on server 110, server 130, or terminal 120, and these steps may be implemented by network 140 according to predetermined logic.
[0038] Optionally, server 110 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The aforementioned terminals 120 and 130 can include, but are not limited to, at least one of the following: mobile phones, tablets, laptops, desktop PCs, digital televisions, and other terminals capable of displaying content. Network 140 can be, for example, a wide area network (WAN), a local area network (LAN), a wireless network, a public telephone network, an intranet, and any other type of network well known to those skilled in the art. It should also be noted that the scenarios described above are merely one example in which embodiments of this disclosure can be implemented and are not restrictive.
[0039] It should be noted that the scenario described above is merely one example in which embodiments of this disclosure can be implemented, and is not restrictive. For example, in some exemplary scenarios, hyperparameter optimization may also be implemented on a specific terminal or server.
[0040] Figure 2 An exemplary flowchart of a hyperparameter optimization method 200 according to some embodiments of this application is shown. Method 200 can be implemented, for example, in applications such as... Figure 1 On server 110, this is not a limitation. Figure 2 As shown, method 200 includes steps 210, 220, 230, 240, 250 and 260.
[0041] In step 210, a candidate hyperparameter set is obtained. The candidate hyperparameter set includes multiple candidate hyperparameters and performance metrics for a first portion of the candidate hyperparameters. The performance metrics are used to indicate the performance of the machine learning model configured with the hyperparameters. In some embodiments, the performance metrics of the first portion of the hyperparameters can be obtained by testing the performance of the machine learning model configured with the first portion of the hyperparameters. In some embodiments, step 210 can be used as an initialization step.
[0042] In some embodiments, obtaining a candidate hyperparameter set may include: obtaining multiple candidate hyperparameters; for example, sampling the range of candidate hyperparameters to obtain multiple candidate hyperparameters; determining a first portion of hyperparameters among the multiple candidate hyperparameters; for example, the first portion of hyperparameters may be determined by random sampling from the candidate hyperparameters or by a specific algorithm; obtaining performance metrics of the first portion of hyperparameters among the multiple candidate hyperparameters; and establishing a candidate hyperparameter set, which includes multiple candidate hyperparameters and performance metrics of the first portion of hyperparameters among the multiple candidate hyperparameters. For example, candidate hyperparameters may be multiple learning rates, and the performance metrics of the first portion of hyperparameters may be the performance of the model after configuring the machine learning model with the first portion of learning rates.
[0043] As an example, determining the first part of hyperparameters from multiple candidate hyperparameters may include: dividing the candidate hyperparameters into a first number of candidate hyperparameter groups, where each candidate hyperparameter group has an equal probability of containing the optimal hyperparameter; and randomly selecting a second number of candidate hyperparameters from each of the first number of candidate hyperparameter groups as the first part of hyperparameters. For example, the cumulative density functions (CDFs) of multiple candidate hyperparameters can be obtained, and then the multiple candidate hyperparameters can be divided into n partitions according to their CDFs, with each partition having an equal probability. Then, the same number of candidate hyperparameters can be selected from each partition, and the selected hyperparameters can be used as the first part of hyperparameters.
[0044] In some embodiments, obtaining the performance metrics of a first portion of hyperparameters from a plurality of candidate hyperparameters may include: obtaining a training set for training a machine learning model, such as obtaining historical data of the target object to be learned; configuring a machine learning model for each of the first portion of hyperparameters, such as configuring the machine learning model with a first portion learning rate; training the machine learning model with the training set to determine the parameters of the machine learning model, such as training a machine learning model with the training set and configured learning rate to obtain a trained machine learning model; testing the performance of the trained machine learning model to determine the performance metrics of the trained machine learning model; and determining the performance metrics of the trained machine learning model as the performance metrics of the corresponding hyperparameters in the first portion of hyperparameters. As an example, when an application scenario requires configuring hyperparameters for a machine learning model used to predict the behavior of a target object, the training set used to train the machine learning model may include historical data of the target object, and then testing the prediction accuracy of the trained machine learning model for the target object as a performance metric of the hyperparameter.
[0045] In step 220, based on the performance metrics of the first portion of hyperparameters among the multiple candidate hyperparameters, the confidence level of each hyperparameter among the non-first portion of the multiple candidate hyperparameters is determined as the optimal hyperparameter, where the optimal hyperparameter refers to the hyperparameter with the best performance metrics. In some embodiments, the confidence level can be obtained using an acquisition function. In some embodiments, step 220 can be used as a confidence level determination step.
[0046] In step 230, the optimal candidate hyperparameters among the non-first part hyperparameters are determined based on the confidence that each hyperparameter in the non-first part is the optimal hyperparameter, and the performance index of the optimal candidate hyperparameters is obtained. In some embodiments, whether the optimal candidate hyperparameter is the optimal hyperparameter is determined based on its performance index. In some embodiments, step 230 can be used as an optimal candidate step.
[0047] In some embodiments, obtaining the performance metrics of the optimal candidate hyperparameters may include: acquiring a training set for training a machine learning model, such as historical data of the target object to be learned; configuring the machine learning model with each of the optimal candidate hyperparameters, such as configuring the machine learning model with the optimal candidate learning rate; training the machine learning model with the training set to determine the parameters of the machine learning model, such as training the machine learning model with the training set and the configured learning rate to obtain the trained machine learning model; testing the performance of the trained machine learning model to determine the performance metrics of the trained machine learning model; and determining the performance metrics of the trained machine learning model as the performance metrics of the corresponding hyperparameters among the optimal candidate hyperparameters. As an example, when an application scenario requires configuring hyperparameters for a machine learning model used to predict the behavior of a target object, the training set used to train the machine learning model may include historical data of the target object, and the prediction accuracy of the trained machine learning model for the target object is then tested as a performance metric for the hyperparameters.
[0048] In step 240, it is determined whether the performance index of the optimal candidate hyperparameter meets the predetermined conditions.
[0049] In some embodiments, the predetermined condition may include: the performance metric of the optimal candidate hyperparameter exceeds a predetermined threshold or the number of times the decision step is executed exceeds a predetermined number. As an example, the number of iterations of the step can be limited to terminate hyperparameter optimization, for instance, limiting the predetermined condition to 100 executions of the decision step.
[0050] In step 250, in response to the fact that the performance index of the optimal candidate hyperparameter does not meet the predetermined conditions, the optimal candidate hyperparameter is added to the first part of the hyperparameters, the performance index of the optimal candidate hyperparameter is added to the candidate hyperparameter set, and then the process proceeds to step 220 for the next iteration.
[0051] In step 260, in response to the performance index of the optimal candidate hyperparameter satisfying a predetermined condition, the optimal candidate hyperparameter is determined as the optimal hyperparameter. In some embodiments, steps 240, 250, and 260 may be part of a determination step.
[0052] As can be seen, Method 200 achieves fast, efficient and accurate hyperparameter optimization by executing the above steps according to predetermined logic.
[0053] Figure 3 An exemplary flowchart of a confidence level determination method 300 according to some embodiments of this application is shown. As an example, method 300 can be performed in step 220 described above. Figure 3 As shown, method 300 includes steps 310, 320, 330 and 340.
[0054] In step 310, a first transformation is performed on the hyperparameters in the hyperparameter set to determine the transformed hyperparameter values, which satisfy a normal distribution. As an example, the first transformation may include the Kumaraswamy transformation, such that the transformed hyperparameters satisfy a normal distribution.
[0055] In some embodiments, performing a first transformation on the hyperparameters in the hyperparameter set to determine the hyperparameter transformation value may include: obtaining a normalized value of the current hyperparameter, the normalized value being between 0 and 1, for example, by performing a normalization operation on the hyperparameter; performing a first mapping on the normalized value of the current hyperparameter to determine a first mapping value of the normalized value of the current hyperparameter; and determining the first mapping value of the normalized value of the current hyperparameter as the hyperparameter transformation value of the current hyperparameter, wherein the normalized value of the current hyperparameter and the first mapping value of the normalized value of the current hyperparameter satisfy the following equation: Among them, h i The normalized value of the hyperparameter, H i The first mapped value of the normalized hyperparameter is denoted as 'a', where 'a' is the normalized value of the smallest hyperparameter in the hyperparameter set, and 'b' is the normalized value of the largest hyperparameter in the hyperparameter set.
[0056] In step 320, a second transformation is performed on the performance indices of the first portion of the hyperparameters in the hyperparameter set to determine the transformed performance index values of the first portion of the hyperparameters. These transformed performance index values satisfy a normal distribution. As an example, the second transformation may include a z-score normalization transformation.
[0057] In some embodiments, performing a second transformation on the performance indices of a first portion of the hyperparameters in the hyperparameter set to determine the transformed performance indices of the first portion of the hyperparameters may include: obtaining the performance indices of the hyperparameters from the hyperparameter set; performing a second mapping on the performance indices of the hyperparameters to determine the second mapped values of the performance indices of the hyperparameters; and determining the second mapped values of the performance indices of the hyperparameters as the transformed performance indices of the hyperparameters, wherein the performance indices of the hyperparameters and the second mapped values of the performance indices of the hyperparameters satisfy the equation: Among them, P i p refers to the second mapping value of the performance index of the hyperparameter. i Hyperparameters refer to performance metrics. The mean of the performance metrics in the set of hyperparameters. The standard deviation of the performance metrics in the set of hyperparameters.
[0058] In step 330, based on the hyperparameter transformation values and performance index transformation values of the first part of hyperparameters, for each hyperparameter outside the first part of hyperparameters, a first confidence level is determined for the hyperparameter transformation value that has the largest performance index transformation value. In some embodiments, the confidence level can be determined using an acquisition function.
[0059] In step 340, the first confidence level of the hyperparameter transformation value of each hyperparameter is determined as the confidence level that each corresponding hyperparameter is the optimal hyperparameter.
[0060] As can be seen, Method 300 performs a first transformation on the hyperparameters and a second transformation on the performance indices of the first part of the hyperparameters to minimize noise interference. Furthermore, based on the transformed hyperparameter values and their performance index transformation values of the first part of the hyperparameters, it determines the confidence level of each hyperparameter outside the first part that has the largest performance index transformation value. This results in a more accurate and robust confidence level. By performing the above steps, Method 300 lays the foundation for achieving fast, efficient, and accurate hyperparameter optimization.
[0061] In some embodiments, step 330, based on the hyperparameter transformation values and performance index transformation values of the first portion of hyperparameters, determining the first confidence level for the hyperparameter transformation value of each hyperparameter in the non-first portion of hyperparameters having the largest performance index transformation value may include: establishing a statistical model based on the hyperparameter transformation values and performance index transformation values of the first portion of hyperparameters, the statistical model being used to determine the expectation and variance of the performance index transformation values of the hyperparameter transformation values; determining the expectation and variance of the performance index transformation values of each hyperparameter in the non-first portion of hyperparameters based on the statistical model; and determining the first confidence level for the hyperparameter transformation value of the corresponding hyperparameter having the largest performance index transformation value based on the expectation and variance of the performance index transformation values of the hyperparameter transformation values of each hyperparameter in the non-first portion of hyperparameters. As an example, the statistical model may predict the performance index of other hyperparameters based on the performance index of known hyperparameters, and provide the expectation and variance of the possible distribution of the performance index. As an example, the statistical model may include one or both of a Gaussian mixture model and a Gaussian model.
[0062] In some embodiments, determining a first confidence level that the hyperparameter transformation value of a corresponding hyperparameter has the largest performance index transformation value, based on the expected value and variance of the performance index transformation value of each hyperparameter transformation value in the non-first part of hyperparameters, includes: obtaining a confidence function, which is used to determine the confidence level based on the expected value and variance; for each hyperparameter transformation value in the non-first part of hyperparameters, substituting the expected value and variance of its performance index transformation value into the confidence function; and determining the output of the confidence function as the first confidence level that the hyperparameter transformation value of the corresponding hyperparameter has the largest performance index transformation value. As an example, the confidence function may include an acquisition function used to determine the confidence level based on the expected value and variance.
[0063] In some embodiments, obtaining the confidence function includes: obtaining a first acquisition function, a second acquisition function, and a third acquisition function; and determining a confidence function based on the first acquisition function, the second acquisition function, and the third acquisition function, wherein the confidence function, the first acquisition function, the second acquisition function, and the third acquisition function satisfy the following equation: in, Confidence function Refers to the first acquisition function. Refers to the second acquisition function. Refers to the third acquisition function. It is the weight of the first acquisition function. It is the weight of the second acquisition function. It is the weight of the third acquisition function and .
[0064] In some embodiments, the first, second, and third acquisition functions each include one or more of the following: an expected increment acquisition function, a probability increment acquisition function, and a confidence upper bound acquisition function; and the weights of the first, second, and third acquisition functions are all between 0 and 1 and are randomly generated. Since the confidence function includes three acquisition functions, the confidence value does not overly depend on any single acquisition function, increasing the generality of the ultimately determined optimal hyperparameters. Furthermore, because the weights of the acquisition functions are randomly generated, the probability of the machine learning model overfitting to certain training sets is reduced, increasing the sparsity of the machine learning model and resulting in better adaptability and accuracy of the obtained optimal hyperparameters.
[0065] Figure 4 An exemplary flowchart 400 of a hyperparameter optimization method according to some embodiments of this application is shown. Figure 4 As shown, method 400 may include steps S410, S420, S430, S440 and S450.
[0066] In step S410, candidate hyperparameters are first obtained, and the Latin hypercube sampling method is used to sample the candidate hyperparameters, and the performance index of the sampled hyperparameters is obtained. For example, the M candidate hyperparameters are divided into n equal probability spaces by the cumulative density function, and then m hyperparameters are randomly selected from each equal probability space. The total of m*n hyperparameters are then used as the sampled hyperparameters, and the performance index of these m*n hyperparameters is obtained.
[0067] In step S420, the Kumaraswamy transform is performed on the M candidate hyperparameters to obtain hyperparameter transformation values, and the z-score transform is performed on the performance indexes of the m*n hyperparameters to obtain performance index transformation values. Then, a Gaussian mixture model is established using the hyperparameter transformation values and performance index transformation values of the m*n hyperparameters to predict the performance index transformation values of the M candidate hyperparameters.
[0068] In step S430, the expected value and variance of the performance index transformation values of the M candidate hyperparameters are obtained based on the established Gaussian mixture model. Then, the optimal candidate hyperparameter among the M candidate hyperparameters is determined using a confidence function that combines the expectation increment function (EI), the probability increment function (PI), and the upper confidence bound function (UCB). For example, the confidence function can be determined by adding weights to the EI, PI, and UCB functions respectively.
[0069] In step S440, it is determined whether the performance index of the optimal candidate hyperparameter meets a predetermined condition. If the predetermined condition is met, the process proceeds to step S450, whereby the optimal candidate hyperparameter is determined as the optimal hyperparameter. If the predetermined condition is not met, the process proceeds to step S420, where the optimal candidate hyperparameter and its performance index are added to the next optimization process, and iteration continues. The predetermined condition may include the value of the performance index being greater than a certain threshold or the number of iterations exceeding an iteration limit. For example, it could be set that if step S440 has been performed 100 times, the iteration can end.
[0070] Figure 5 Exemplary schematic diagrams of a hyperparameter optimization system 500 according to some embodiments of this application are shown. The hyperparameter optimization system 500 is used for online hyperparameter tuning to determine the optimal hyperparameters in real time. It should be noted that in other embodiments, offline hyperparameter tuning can also be performed to determine the optimal hyperparameters offline. For example... Figure 5 The hyperparameter optimization system 500 includes a parameter tuning section and a business section. The business section includes a business backend 510 and a database 520, while the parameter tuning section includes a scheduler 530, a parameter tuning backend 540, and a management backend 550.
[0071] First, the business backend 510 notifies the management backend 550 to initiate a hyperparameter tuning task. The management backend 550 then takes the machine learning model based on the old hyperparameters offline from the business backend 510 and uploads a machine learning model based on the new hyperparameters. Optionally, the new hyperparameters can be the optimal candidate hyperparameters mentioned earlier. Then, the business backend 510 deploys the machine learning model based on the new hyperparameters to multiple application scenarios and collects real-time data (e.g., performance metrics) of the machine learning model in each scenario, storing it in database 520. Next, the scheduler 530 schedules a portion of the data in database 520 according to the needs of hyperparameter tuning and sends it to the hyperparameter tuning backend 510. After obtaining the data from the machine learning model, the hyperparameter tuning backend 540 can determine the optimal candidate hyperparameters based on the embodiments shown in methods 200 and 400 above, and train a new machine learning model based on this, sending it to the management backend 550. The management backend 550 then uploads the latest machine learning model to the business backend 510 and takes the old machine learning model offline. Optionally, the iteration can be terminated by setting the number of times the latest machine learning model is uploaded to the service backend 510 by the management backend 550. In some embodiments, the scheduler 530 can divide the scheduled data into multiple groups and transmit them to the parameter tuning backend separately to enable subsequent experimental control.
[0072] As an example, the hyperparameter optimization system 500 can be used for hyperparameter tuning experiments of deep learning models, such as tuning the learning rate of a deep learning model to obtain the optimal learning rate. For instance, when preparing to determine the optimal learning rate from the range (0, 0.2), this can be achieved through system 500.
[0073] First, the business backend 510 will create a hyperparameter tuning experiment, setting the learning rate range to be searched (0.001, 0.2), the number of groups to be searched in each batch to 3, and the optimization metric to the average viewing time per user. After the hyperparameter tuning begins, the business backend 510 will divide all users into six equal parts, each part corresponding to a hyperparameter. Two hyperparameters can be set as a pair, with each pair containing a control version and an experimental version. The learning rate of the control version can be set to increase by 0.001 each time, and the learning rate of the experimental version is configured as the optimal candidate hyperparameter. Then, every 72 hours, the scheduler 530 will send the collected learning rates of the six parts and the corresponding average viewing time per user to the hyperparameter tuning backend 540. The hyperparameter tuning backend will calculate the three learning rates to be searched in the next batch (i.e., the next batch of control group and the next batch of optimal candidate hyperparameters) based on the average viewing time per user and the learning rate range to be searched. Then, the parameter tuning backend 540 sends it to the management backend 550. The management backend 550 takes the six deep learning models corresponding to the three pairs of hyperparameters currently online offline and brings online the six deep learning models corresponding to the next batch of three sets of parameters to be searched. The above steps are repeated until the average viewing time per user that satisfies the business is obtained. The corresponding learning rate is the final optimal learning rate (i.e., the optimal hyperparameter).
[0074] To test the performance of the various hyperparameter optimization methods presented in this disclosure, a first set of datasets and a second set of datasets were randomly obtained and used for the first and second sets of experiments, respectively. In each set of experiments, 10 rounds of testing were conducted, with 100 iterations in each round to obtain the optimal hyperparameters. Finally, the machine learning model configured with the optimal hyperparameters was scored, and the model score with the optimal hyperparameters was used as the performance metric. To more intuitively represent the comparison of the effects of each method, the scores of each model were normalized. Figure 6 , Figure 7 and Figure 8 Schematic diagrams illustrating the performance of hyperparameter optimization methods according to some embodiments of this application are shown.
[0075] like Figure 7 and Figure 8 As shown in the figure, the horizontal axis represents the test sequence number (i.e., the sequence number in 10 rounds of testing, for example, 3 represents the third round of testing), and the vertical axis represents the performance index of the hyperparameters (i.e., the normalized value of the model score corresponding to the hyperparameters). Here, random search means randomly searching for the optimal hyperparameter among the candidate hyperparameters; PI represents the hyperparameter optimization method using the PI sampling function as the confidence function; EI represents the hyperparameter optimization method using the EI sampling function as the confidence function; UCB represents the hyperparameter optimization method using the UCB sampling function as the confidence function; and method 400 represents the hyperparameter optimization method shown in the embodiment of method 400 described above.
[0076] from Figure 7 and Figure 8 It can be seen that, in both sets of experiments, the optimal hyperparameters determined by method 400 exhibit the most stability and the highest performance index. While the random search algorithm can occasionally achieve high performance indexes, its performance is highly unstable. Furthermore, the PI, UCB, and EI algorithms proposed in this disclosure demonstrate better stability and performance indexes than the random search algorithm, but worse than method 400. In some experiments, some algorithms even showed performance indices of 0. Therefore, in terms of demonstrated stability, the method of this disclosure is the best.
[0077] Figure 8 The performance of hyperparameter optimization methods according to some embodiments of this application is illustrated. The average score represents the average score of these algorithms in the first and second sets of experiments. Combined with... Figure 8 It can be seen that the method 400 proposed in this disclosure for determining the optimal hyperparameters has good performance in both accuracy and stability.
[0078] Figure 9 An exemplary structural block diagram of a hyperparameter optimization apparatus 900 according to some embodiments of this application is shown. Figure 9 As shown, the hyperparameter optimization device 900 includes an initialization module 910, a confidence module 920, an optimal candidate module 930, and a judgment module 940. The hyperparameter optimization device 900 is configured to perform an initialization step, and then perform a confidence determination step, an optimal candidate step, and a judgment step according to predetermined logic.
[0079] The initialization module 910 is configured to perform an initialization step: obtaining a set of candidate hyperparameters, which includes multiple candidate hyperparameters and performance metrics for a first subset of the candidate hyperparameters. These performance metrics indicate the performance of a machine learning model configured with the corresponding hyperparameters. In some embodiments, the performance metrics for the first subset of hyperparameters can be obtained by testing the performance of a machine learning model configured with the first subset of hyperparameters. For example, candidate hyperparameters may be multiple learning rates, and the performance metrics for the first subset of hyperparameters may be the performance of the model after configuring it with the first subset of learning rates.
[0080] The confidence module 920 is configured to perform a confidence determination step: based on the performance metrics of a first portion of the candidate hyperparameters, for each hyperparameter not in the first portion of the candidate hyperparameters, determine its confidence level as the optimal hyperparameter, where the optimal hyperparameter is the one with the best performance metrics. In some embodiments, the confidence level may include obtaining it using an acquisition function. In some embodiments, step 220 may serve as the confidence determination step.
[0081] The optimal candidate module 930 is configured to perform the optimal candidate step: determining the optimal candidate hyperparameter among the non-first part hyperparameters based on the confidence that each hyperparameter in the non-first part is the optimal hyperparameter, and obtaining the performance index of the optimal candidate hyperparameter. In some embodiments, whether the optimal candidate hyperparameter is the optimal hyperparameter is determined based on its performance index.
[0082] The judgment module 940 is configured to perform the following judgment steps: determining whether the performance metric of the optimal candidate hyperparameter meets a predetermined condition; if the performance metric of the optimal candidate hyperparameter does not meet the predetermined condition, adding the optimal candidate hyperparameter to the first part of hyperparameters and adding the performance metric of the optimal candidate hyperparameter to the candidate hyperparameter set, proceeding to the confidence determination step; and if the performance metric of the optimal candidate hyperparameter meets the predetermined condition, determining the optimal candidate hyperparameter as the optimal hyperparameter. In some embodiments, the predetermined condition may include: the performance metric of the optimal candidate hyperparameter exceeds a predetermined threshold or the number of times the judgment step is executed exceeds a predetermined number. As an example, the number of iterations of the step can be limited to end the hyperparameter optimization, for example, limiting the predetermined condition to 100 executions of the judgment step.
[0083] Optionally, the hyperparameter optimization device 900 can be used to execute the above-described methods 200, 300, 400, etc. It is evident that by using the hyperparameter optimization device 900, the above steps can be executed according to predetermined logic to achieve fast, efficient, and accurate hyperparameter optimization.
[0084] Figure 10 The illustration depicts an example system 1000, which includes an example computing device 1010 representing one or more systems and / or devices that can implement the various technologies described herein. The computing device 1010 may be, for example, a server of a service provider, a device associated with a server, a system-on-a-chip, and / or any other suitable computing device or computing system. (Refer to above) Figure 9 The described hyperparameter optimization apparatus 900 can take the form of a computing device 1010. Alternatively, the hyperparameter optimization apparatus 900 can be implemented as a computer program as an application 1016.
[0085] The example computing device 1010 shown includes a processing system 1011, one or more computer-readable media 1012, and one or more I / O interfaces 1013, all communicatively coupled to each other. Although not shown, the computing device 1010 may also include a system bus or other data and command transfer system that couples the various components to each other. The system bus may include any or a combination of different bus architectures, such as a memory bus or memory controller, a peripheral bus, a universal serial bus, and / or a processor or local bus utilizing any of the various bus architectures. Various other examples, such as control and data lines, are also conceived.
[0086] Processing system 1011 represents the functionality of performing one or more operations using hardware. Therefore, processing system 1011 is illustrated as including hardware elements 1014 that can be configured as processors, function blocks, etc. This may include application-specific integrated circuits (ASICs) or other logic devices formed using one or more semiconductors in the hardware. Hardware element 1014 is not limited by the materials in which it is formed or the processing mechanism employed therein. For example, a processor may consist of semiconductors and / or transistors (e.g., integrated circuits (ICs)). In such a context, processor-executable instructions may be electronically executable instructions.
[0087] Computer-readable medium 1012 is illustrated as including memory / storage device 1015. Memory / storage device 1015 represents a memory / storage capacity associated with one or more computer-readable media. Memory / storage device 1015 may include volatile media (such as random access memory (RAM)) and / or non-volatile media (such as read-only memory (ROM), flash memory, optical disk, magnetic disk, etc.). Memory / storage device 1015 may include fixed media (e.g., RAM, ROM, fixed hard disk drive, etc.) and removable media (e.g., flash memory, removable hard disk drive, optical disk, etc.). Computer-readable medium 1012 may be configured in various other ways as further described below.
[0088] One or more I / O interfaces 1013 represent functions that allow users to input commands and information to the computing device 1010 using various input devices and optionally also allow information to be presented to the user and / or other components or devices using various output devices. Examples of input devices include keyboards, cursor control devices (e.g., mice), microphones (e.g., for voice input), scanners, touch functions (e.g., capacitive or other sensors configured to detect physical touch), cameras (e.g., capable of detecting non-touch-related movements as gestures using visible or invisible wavelengths (such as infrared frequencies), etc. Examples of output devices include display devices, speakers, printers, network interface cards, haptic-responsive devices, etc. Therefore, the computing device 1010 can be configured to support user interaction in various ways as further described below.
[0089] The computing device 1010 also includes an application 1016. The application 1016 may be, for example, a software instance of a hyperparameter optimization device 900, and may implement the techniques described herein in combination with other elements in the computing device 1010.
[0090] This document describes various technologies within the general context of software and hardware components or program modules. Generally, these modules include routines, programs, objects, elements, components, data structures, etc., that perform specific tasks or implement specific abstract data types. As used herein, the terms "module," "function," and "component" generally refer to software, firmware, hardware, or a combination thereof. The technologies described herein are characterized as platform-independent, meaning that these technologies can be implemented on a variety of computing platforms with various processors.
[0091] Implementations of the described modules and technologies may be stored on or transmitted across some form of computer-readable medium. The computer-readable medium may include a variety of media accessible by the computing device 1010. By way of example and not limitation, the computer-readable medium may include "computer-readable storage media" and "computer-readable signal media".
[0092] In contrast to simple signal transmission, carrier waves, or signals themselves, a "computer-readable storage medium" refers to a medium and / or device capable of persistently storing information, and / or a tangible storage device. Therefore, a computer-readable storage medium refers to a non-signal-bearing medium. Computer-readable storage media include hardware such as volatile and non-volatile, removable and non-removable media and / or storage devices implemented using methods or techniques suitable for storing information (such as computer-readable instructions, data structures, program modules, logic elements / circuits, or other data). Examples of computer-readable storage media may include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical storage devices, hard disks, cassette tapes, magnetic tapes, disk storage devices or other magnetic storage devices, or other storage devices, tangible media, or articles of art suitable for storing desired information and accessible by a computer.
[0093] "Computer-readable signal medium" refers to a signal-bearing medium configured to transmit instructions, such as via a network, to computing device 1010. A signal medium typically embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal, such as a carrier wave, data signal, or other transmission mechanism. Signal media also include any information transmission medium. The term "modulated data signal" refers to a signal in which one or more of its characteristics are set or altered to encode information. By way of example and not limitation, communication media include wired media such as wired networks or direct connections, and wireless media such as acoustic, RF, infrared, and other wireless media.
[0094] As previously described, hardware element 1014 and computer-readable medium 1012 represent instructions, modules, programmable device logic, and / or fixed device logic implemented in hardware, which in some embodiments can be used to implement at least some aspects of the techniques described herein. Hardware elements may include components of integrated circuits or systems-on-a-chip, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), complex programmable logic devices (CPLDs), and other implementations or other hardware devices in silicon. In this context, hardware elements can serve as processing devices for executing program tasks defined by instructions, modules, and / or logic embodied by the hardware element, and as hardware devices for storing instructions for execution, such as the previously described computer-readable storage medium.
[0095] The foregoing combinations can also be used to implement the various techniques and modules described herein. Therefore, software, hardware, or program modules and other program modules can be implemented as one or more instructions and / or logic embodied on some form of computer-readable storage medium and / or by one or more hardware elements 1014. The computing device 1010 can be configured to implement specific instructions and / or functions corresponding to the software and / or hardware modules. Thus, for example, by using the computer-readable storage medium and / or hardware elements 1014 of a processing system, modules can be implemented at least partially in hardware as modules executable as software by the computing device 1010. Instructions and / or functions can be executable / operable by one or more articles of art (e.g., one or more computing devices 1010 and / or processing systems 1011) to implement the techniques, modules, and examples described herein.
[0096] In various embodiments, the computing device 1010 can be configured in various ways. For example, the computing device 1010 can be implemented as a computer-type device, including personal computers, desktop computers, multi-screen computers, laptop computers, netbooks, etc. The computing device 1010 can also be implemented as a mobile device, including mobile devices such as mobile phones, portable music players, portable gaming devices, tablet computers, multi-screen computers, etc. The computing device 1010 can also be implemented as a television-type device, including devices with or connected to a generally large screen in a leisure viewing environment. These devices include televisions, set-top boxes, game consoles, etc.
[0097] The techniques described herein can be supported by these various configurations of computing device 1010, and are not limited to specific examples of the techniques described herein. Functionality can also be implemented, wholly or partially, on the “cloud” 1020 using distributed systems, such as through platform 1022 as described below.
[0098] Cloud 1020 includes and / or represents platform 1022 for resource 1024. Platform 1022 abstracts the underlying functionality of the hardware (e.g., server) and software resources of cloud 1020. Resource 1024 may include applications and / or data that can be used when performing computer processing on a server located remotely from computing device 1010. Resource 1024 may also include services provided via the Internet and / or via subscriber networks such as cellular or Wi-Fi networks.
[0099] Platform 1022 can abstract resources and functions to connect computing device 1010 to other computing devices. Platform 1022 can also be used to abstract resource hierarchy to provide a corresponding level of hierarchy for any encountered needs for resource 1024 implemented via platform 1022. Therefore, in interconnect device embodiments, the implementation of the functions described herein can be distributed throughout system 1000. For example, functions can be implemented partly on computing device 1010 and partly through platform 1022, which abstracts the functions of cloud 1020.
[0100] This disclosure provides a computer-readable storage medium having computer-readable instructions stored thereon, which, when executed, implement the above-described method for hyperparameter optimization.
[0101] This disclosure provides a computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computing device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computing device to perform the hyperparameter optimization methods provided in the various alternative implementations described above.
[0102] It should be understood that, for clarity, embodiments of this disclosure have been described with reference to different functional units. However, it will be apparent that, without departing from this disclosure, the functionality of each functional unit may be implemented in a single unit, in multiple units, or as part of other functional units. For example, functionality described as being performed by a single unit may be performed by multiple different units. Therefore, references to a particular functional unit are considered merely as references to the appropriate unit used to provide the described functionality, and not as indicating a strict logical or physical structure or organization. Thus, this disclosure may be implemented in a single unit, or may be physically and functionally distributed among different units and circuits.
[0103] It will be understood that although the terms first, second, third, etc., may be used herein to describe various devices, elements, components, or parts, these devices, elements, components, or parts should not be limited by these terms. These terms are used only to distinguish one device, element, component, or part from another device, element, component, or part.
[0104] Although this disclosure has been described in conjunction with some embodiments, it is not intended to be limited to the specific forms set forth herein. Rather, the scope of this disclosure is limited only by the appended claims. Additionally, although individual features may be included in different claims, these may be advantageously combined, and inclusion in different claims does not imply that such a combination of features is not feasible and / or advantageous. The order of features in the claims does not imply that the features must be in any particular order of their operation. Furthermore, in the claims, the word "comprising" does not exclude other elements, and the terms "a" or "an" do not exclude a plurality. Reference numerals in the claims are provided only by way of explicit example and should not be construed as limiting the scope of the claims in any way.
[0105] It is understood that the specific embodiments of this application involve user login data such as usernames and passwords. When the embodiments of this application involving such data are applied to specific products or technologies, user permission or consent is required, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions.
Claims
1. A hyperparameter optimization method for a machine learning model, characterized in that, The machine learning model is used for content presentation, and the method includes performing an initialization step, and executing a confidence determination step, an optimal candidate step, and a judgment step according to predetermined logic: The initialization step is as follows: Obtain a candidate hyperparameter set, which includes multiple candidate hyperparameters and a performance index of the first part of the hyperparameters among the multiple candidate hyperparameters. The performance index is used to indicate the average viewing time per user when the model configured with the corresponding hyperparameters for content presentation is presented to the user. The confidence determination step is as follows: based on the performance index of the first part of the multiple candidate hyperparameters, for each hyperparameter that is not in the first part of the multiple candidate hyperparameters, determine the confidence that it is the optimal hyperparameter, where the optimal hyperparameter refers to the hyperparameter with the highest average viewing time per person. The optimal candidate step involves: determining the optimal candidate hyperparameter among the non-first part hyperparameters based on the confidence that each hyperparameter in the non-first part is the optimal hyperparameter, and obtaining the performance index of the optimal candidate hyperparameter; and... The judgment step involves determining whether the performance index of the optimal candidate hyperparameter meets predetermined conditions. If the performance metric of the optimal candidate hyperparameter does not meet the predetermined condition, the optimal candidate hyperparameter is added to the first part of hyperparameters, and the performance metric of the optimal candidate hyperparameter is added to the candidate hyperparameter set, then proceeding to the confidence determination step. If the performance index of the optimal candidate hyperparameter satisfies the predetermined condition, the optimal candidate hyperparameter is determined as the optimal hyperparameter. The hyperparameter is the learning rate. The confidence level determination step further includes: A first transformation is performed on the hyperparameters in the hyperparameter set to determine the hyperparameter transformation values, wherein the hyperparameter transformation values satisfy a normal distribution; A second transformation is performed on the performance index of the first part of the hyperparameter set to determine the transformed performance index value of the first part of the hyperparameter set, wherein the transformed performance index value satisfies a normal distribution. Based on the hyperparameter transformation values and performance index transformation values of the first portion of hyperparameters, for each hyperparameter in the non-first portion of hyperparameters, a first confidence level is determined for its maximum performance index transformation value; and, The first confidence level of the hyperparameter transformation value of each hyperparameter is determined as the confidence level that the corresponding hyperparameter is the optimal hyperparameter.
2. The method according to claim 1, wherein performing a first transformation on the hyperparameters in the hyperparameter set to determine the hyperparameter transformation values comprises: Obtain the normalized value of the current hyperparameter, where the normalized value is between 0 and 1; Perform a first mapping on the normalized value of the current hyperparameter to determine the first mapping value of the normalized value of the current hyperparameter; as well as, The first mapping value of the normalized value of the current hyperparameter is determined as the hyperparameter transformation value of the current hyperparameter. The normalized value of the current hyperparameter and the first mapping value of the normalized value of the current hyperparameter satisfy the following equation: wherein h i is the normalized value of the current hyperparameter, H i is the first mapping value of the normalized value of the current hyperparameter, a is the normalized value of the smallest hyperparameter in the set of hyperparameters, and b is the normalized value of the largest hyperparameter in the set of hyperparameters.
3. The method according to claim 1, wherein performing a second transformation on the performance index of the first portion of the hyperparameters of the hyperparameter set to determine the transformed performance index value of the first portion of the hyperparameters includes: Obtain the performance metrics of the hyperparameters from the set of hyperparameters; A second mapping is performed on the performance index of the hyperparameter to determine the second mapping value of the performance index of the hyperparameter. as well as, The second mapping value of the performance index of the hyperparameter is determined as the transformed value of the performance index of the hyperparameter. The performance index of the hyperparameter and the second mapping value of the performance index of the hyperparameter satisfy the following equation: Among them, P i It is the second mapping value of the performance index of the hyperparameter, p i These are the performance metrics of the hyperparameters. It is the mean of the performance metrics in the set of hyperparameters. It is the standard deviation of the performance metrics in the set of hyperparameters.
4. The method according to claim 1, wherein, based on the hyperparameter transformation values and performance index transformation values of the first portion of hyperparameters, determining the first confidence level of the hyperparameter transformation value with the largest performance index transformation value for each of the non-first portion of hyperparameters includes: A statistical model is established based on the hyperparameter transformation values and performance index transformation values of the first part of hyperparameters. The statistical model is used to determine the expectation and variance of the performance index transformation values of the hyperparameter transformation values. Based on the statistical model, for each hyperparameter in the non-first part of the hyperparameters, determine the expected value and variance of its performance index transformation value; as well as, Based on the expectation and variance of the performance index transformation value of the hyperparameter transformation value of each hyperparameter in the non-first part of the hyperparameters, a first confidence level is determined for the hyperparameter transformation value of the corresponding hyperparameter to have the maximum performance index transformation value.
5. The method according to claim 4, wherein determining the first confidence level that the hyperparameter transformation value of the corresponding hyperparameter has the largest performance index transformation value based on the expectation and variance of the performance index transformation value of the hyperparameter transformation value of each hyperparameter in the non-first part of hyperparameters includes: Obtain a confidence function, which is used to determine the confidence level based on expectation and variance; For each hyperparameter in the non-first part of the hyperparameters, the expected value and variance of its performance index transformation value are substituted into the confidence function; as well as, The output of the confidence function is determined as the first confidence level of the hyperparameter transformation value with the maximum performance index transformation value.
6. The method according to claim 5, wherein obtaining the confidence function comprises: Obtain the first acquisition function, the second acquisition function, and the third acquisition function; as well as, The confidence function is determined based on the first acquisition function, the second acquisition function, and the third acquisition function, wherein the confidence function, the first acquisition function, the second acquisition function, and the third acquisition function satisfy the following equation: in, It is the confidence function. It is the first acquisition function. It is the second acquisition function. It is the third acquisition function. It is the weight of the first acquisition function. It is the weight of the second acquisition function. Is the weight of the third acquisition function and .
7. The method according to claim 6, wherein the first acquisition function, the second acquisition function, and the third acquisition function each comprise: The expected incremental acquisition function, the probability incremental acquisition function, and the confidence upper bound acquisition function are one or more of the following: the weights of the first acquisition function, the second acquisition function, and the third acquisition function are all between 0 and 1 and are randomly generated.
8. The method according to claim 4, wherein the statistical model includes one or both of a Gaussian mixture model and a Gaussian model.
9. The method according to claim 1, wherein obtaining the candidate hyperparameter set comprises: Obtain multiple candidate hyperparameters; Determine the first portion of hyperparameters among the plurality of candidate hyperparameters; Obtain the performance metrics of the first portion of hyperparameters from the plurality of candidate hyperparameters; as well as, Establish the candidate hyperparameter set, which includes the plurality of candidate hyperparameters and the performance metrics of the first portion of the candidate hyperparameters.
10. The method of claim 9, wherein determining the first portion of the hyperparameters among the plurality of candidate hyperparameters comprises: The candidate hyperparameters are divided into a first number of candidate hyperparameter groups, and each candidate hyperparameter group in the first number of candidate hyperparameter groups has an equal probability of containing the optimal hyperparameter. as well as, A second number of candidate hyperparameters are randomly selected from each of the first number of candidate hyperparameter groups as the first part of the hyperparameters.
11. The method according to claim 9, wherein obtaining the performance index of the first portion of the hyperparameters among the plurality of candidate hyperparameters includes: Obtain a training set for training the machine learning model; For each of the hyperparameters in the first part, configure the machine learning model accordingly; The machine learning model is trained using the training set to determine the parameters of the machine learning model; Test the performance of the trained machine learning model to determine the performance metrics of the trained machine learning model; as well as, The performance metrics of the trained machine learning model are determined as the performance metrics of the corresponding hyperparameters in the first part of the hyperparameters.
12. The method according to claim 1, wherein the predetermined conditions include: The performance index of the optimal candidate hyperparameter exceeds a predetermined threshold or the number of times the judgment step is executed exceeds a predetermined number.
13. The method according to claim 1, wherein obtaining the performance index of the optimal candidate hyperparameter includes: Obtain a training set for training the machine learning model; For each of the optimal candidate hyperparameters, configure the machine learning model accordingly; The machine learning model is trained using the training set to determine the parameters of the machine learning model; Test the performance of the trained machine learning model to determine the performance metrics of the trained machine learning model; as well as, The performance metrics of the trained machine learning model are determined as the performance metrics of the corresponding hyperparameters among the optimal candidate hyperparameters.
14. A hyperparameter optimization device for a machine learning model, characterized in that, The machine learning model is used for content presentation. The hyperparameter optimization device is configured to perform an initialization step, and execute a confidence determination step, an optimal candidate step, and a judgment step according to predetermined logic. The hyperparameter optimization device includes: An initialization module is configured to perform the initialization steps: obtaining a candidate hyperparameter set, the candidate hyperparameter set including multiple candidate hyperparameters and a performance index of a first portion of the hyperparameters among the multiple candidate hyperparameters, the performance index being used to indicate the average viewing time per user when the model configured with the corresponding hyperparameters for content presentation is presented to the user; The confidence module is configured to perform the confidence determination step: based on the performance index of the first part of the candidate hyperparameters, for each hyperparameter that is not in the first part of the candidate hyperparameters, determine the confidence that it is the optimal hyperparameter, wherein the optimal hyperparameter refers to the hyperparameter with the highest average viewing time per person; An optimal candidate module is configured to perform the optimal candidate step: determining the optimal candidate hyperparameter among the non-first part hyperparameters based on the confidence that each hyperparameter in the non-first part is the optimal hyperparameter, and obtaining the performance index of the optimal candidate hyperparameter; and, The judgment module is configured to perform the judgment step: determining whether the performance index of the optimal candidate hyperparameter meets predetermined conditions. If the performance metric of the optimal candidate hyperparameter does not meet the predetermined condition, the optimal candidate hyperparameter is added to the first part of hyperparameters, and the performance metric of the optimal candidate hyperparameter is added to the candidate hyperparameter set, then proceeding to the confidence determination step. If the performance index of the optimal candidate hyperparameter satisfies the predetermined condition, the optimal candidate hyperparameter is determined as the optimal hyperparameter. The hyperparameter is the learning rate. The confidence level determination step further includes: A first transformation is performed on the hyperparameters in the hyperparameter set to determine the hyperparameter transformation values, wherein the hyperparameter transformation values satisfy a normal distribution; A second transformation is performed on the performance index of the first part of the hyperparameter set to determine the transformed performance index value of the first part of the hyperparameter set, wherein the transformed performance index value satisfies a normal distribution. Based on the hyperparameter transformation values and performance index transformation values of the first portion of hyperparameters, for each hyperparameter in the non-first portion of hyperparameters, a first confidence level is determined for its maximum performance index transformation value; and, The first confidence level of the hyperparameter transformation value of each hyperparameter is determined as the confidence level that the corresponding hyperparameter is the optimal hyperparameter.
15. A computing device, comprising: Memory, which is configured to store computer-executable instructions; as well as A processor configured to perform the method according to any one of claims 1-13 when the computer-executable instructions are executed by the processor.
16. A computer-readable storage medium storing computer-executable instructions that, when executed, implement the method according to any one of claims 1-13.
17. A computer program product comprising a computer program that, when executed by a processor, implements the steps of the method according to any one of claims 1 to 13.
Citation Information
Patent Citations
Using Metamodeling for Fast and Accurate Hyperparameter optimization of Machine Learning and Deep Learning Models
US20200380378A1
Methods, apparatus, and articles of manufacture to improve automated machine learning
US20210117841A1