Hyperparameter Processing Method, Apparatus, Electronic Device, and Storage Medium
By interpolation of the hyperparameters that have not returned the verification results and generating the hyperparameters to be verified, the problems of exploratory decline and resource waste in parallel hyperparameter optimization are solved, and verification efficiency and resource utilization are improved.
Patent Information
- Application Number
- CN202111176055.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-09
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2041-10-09
AI Technical Summary
During the parallel hyperparameter optimization and verification process, hyperparameters that did not return the verification results exist, resulting in the optimization algorithm's exploration of the hyperparameter space and the waste of verification resources.
By obtaining historical verification results, interpolate the verification results of the target hyperparameters, generate the hyperparameters to be verified, and send them to the queue to obtain the verification results, thereby improving the exploration of the hyperparameter space and avoiding resource waste.
The efficiency and resource utilization of hyperparameter verification are improved, the waste caused by the verification of the same or similar hyperparameters to be verified is reduced, and the exploration of the hyperparameter space is enhanced.
Smart Images

Figure CN114065943B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of machine learning, and in particular, to a method, apparatus, electronic device, and storage medium for hyperparameter processing. Background Art
[0002] When using a machine learning model, a series of hyperparameters need to be specified, and these hyperparameters greatly affect the final effect of the machine learning model. Usually, machine learning application developers need to optimize the hyperparameters and select a set of optimal hyperparameters for the machine learning algorithm to improve the learning performance and effect. For example, taking the learning rate of the model as an example, an overly large learning rate may result in poor convergence of the model, while an overly small learning rate may cause the model to converge slowly and get stuck in a local optimum. To obtain a model with excellent performance, it often takes a lot of time to try various hyperparameter configurations to achieve better results.
[0003] In the related art, during the process of parallel hyperparameter optimization and verification, when giving new hyperparameter recommendations, there may be hyperparameters for which the verification is still running and the verification results have not been returned. The existence of hyperparameters for which the verification results have not been returned may lead to a decline in the parallel recommendation effect of the optimization algorithm. This is because when the optimization algorithm makes recommendations based on the verification historical results, if the returned verification historical results remain unchanged and the optimization algorithm is required to give multiple sets of hyperparameters, these hyperparameters may be the same or very similar, resulting in a decrease in the exploration of the hyperparameter space by the optimization algorithm, and the repeated verification of the same or very similar hyperparameters will cause waste of verification resources. Summary of the Invention
[0004] The present disclosure provides a method, apparatus, electronic device, and storage medium for hyperparameter processing to at least solve the problems of decreased exploration of the hyperparameter space and wasted verification resources in the related art. The technical solution of the present disclosure is as follows:
[0005] According to a first aspect of an embodiment of the present disclosure, there is provided a hyperparameter processing method, including:
[0006] Obtaining historical verification results corresponding to historical hyperparameters; the historical verification results are used to characterize the performance data of the trained model corresponding to the historical hyperparameters; the trained model corresponding to the historical hyperparameters is obtained by training a model to be trained based on the historical hyperparameters;
[0007] Performing verification result interpolation on target hyperparameters for which verification results have not been returned based on the historical verification results to obtain interpolation results of the target hyperparameters;
[0008] Generating hyperparameters to be verified based on the historical verification results and the interpolation results of the target hyperparameters;
[0009] Send the hyperparameter to be verified to the queue;
[0010] Obtain the verification result of the hyperparameter to be verified from the queue; the verification result of the hyperparameter to be verified is obtained by multiple verification nodes in the distributed cluster obtaining the hyperparameter to be verified from the queue and verifying the obtained hyperparameter to be verified.
[0011] In an exemplary embodiment, the interpolating the verification result of the target hyperparameter based on the historical verification result to obtain the interpolation result of the target hyperparameter includes:
[0012] Calculate the verification result calculation value for the historical verification result corresponding to the historical hyperparameter;
[0013] Determine the verification result calculation value as the interpolation result of the target hyperparameter.
[0014] In an exemplary embodiment, the historical verification result includes the full - resource verification result obtained by verifying the historical verified hyperparameter based on all resources, and the partial - resource verification result obtained by verifying the historical hyperparameter based on partial resources;
[0015] The calculating the verification result calculation value for the historical verification result corresponding to the historical hyperparameter includes:
[0016] Determine the resource amount of the target hyperparameter; the resource amount is used to characterize the resource quantity information based on which a hyperparameter is verified;
[0017] When the resource amount is all resources, calculate the full - resource verification result to obtain a first calculation value;
[0018] When the resource amount is partial resources, calculate the partial - resource verification result to obtain a second calculation value;
[0019] Determine the first calculation value or the second calculation value as the verification result calculation value.
[0020] In an exemplary embodiment, the generating the hyperparameter to be verified based on the historical verification result and the interpolation result of the target hyperparameter includes:
[0021] Update the current integrated probability proxy model based on the historical verification result and the interpolation result of the target hyperparameter to obtain an updated integrated probability proxy model;
[0022] Generate the hyperparameter to be verified based on the updated integrated probability proxy model.
[0023] In an exemplary embodiment, the method further includes:
[0024] Obtain an initial full - scale resource proxy model and at least one initial partial - scale resource proxy model;
[0025] Assign weights to the initial full - scale resource proxy model and the at least one initial partial - scale resource proxy model; wherein the sum of the weights corresponding to the initial full - scale resource proxy model and the weights corresponding to the at least one initial partial - scale resource proxy model is 1;
[0026] Generate the integrated probability proxy model based on the initial full - scale resource proxy model, the weight of the initial full - scale resource proxy model, the at least one initial partial - scale resource proxy model, and the weights of the at least one initial partial - scale resource proxy model.
[0027] In an exemplary embodiment, the updating the current integrated probability proxy model based on the historical verification result and the imputation result of the target hyperparameter to obtain an updated integrated probability proxy model includes:
[0028] When the imputation result of the target hyperparameter includes a full - scale resource imputation result corresponding to the full - scale resource, update the current full - scale resource proxy model in the current integrated probability proxy model with the full - scale resource imputation result;
[0029] When the imputation result of the target hyperparameter includes a partial - scale resource imputation result corresponding to the partial - scale resource, update the current partial - scale resource proxy model in the current integrated probability proxy model with the partial - scale resource imputation result;
[0030] Obtain the updated integrated probability proxy model based on the updated current full - scale resource proxy model and the updated current partial - scale resource proxy model.
[0031] In an exemplary embodiment, the method further includes:
[0032] Select at least one set of hyperparameter combinations; the hyperparameter combinations include at least two hyperparameters;
[0033] Determine a first verification result corresponding to the at least one set of hyperparameter combinations; the first verification result is a result obtained by verifying the at least one set of hyperparameter combinations based on the full - scale resource;
[0034] Determine a second verification result corresponding to the at least one set of hyperparameter combinations; the second verification result is a result obtained by verifying the at least one set of hyperparameter combinations based on the partial - scale resource;
[0035] Perform a consistency check on the partial order relationship of the first verification result and the partial order relationship of the second verification result;
[0036] Adjust the weights of the full - scale resource proxy model and the at least one partial - scale resource proxy model based on the consistency check result.
[0037] In an exemplary embodiment, the method further includes:
[0038] Based on the weight of the current full - scale resource proxy model and the resource amount corresponding to the current full - scale resource proxy model, obtain a first resource amount;
[0039] Based on the weight of the current partial - scale resource proxy model and the resource amount corresponding to the current partial - scale resource proxy model, obtain a second resource amount;
[0040] According to the first resource amount and the second resource amount, calculate the initial resource amount used when verifying the hyperparameter to be verified.
[0041] According to a second aspect of the embodiments of the present disclosure, there is provided a hyperparameter processing device, including:
[0042] A historical verification result acquisition unit, configured to acquire historical verification results corresponding to historical hyperparameters; the historical verification results are used to characterize the performance data of the trained model corresponding to the historical hyperparameters; the trained model corresponding to the historical hyperparameters is obtained by training a model to be trained based on the historical hyperparameters;
[0043] A result interpolation unit, configured to perform result interpolation on a target hyperparameter for which a verification result has not been returned based on the historical verification results to obtain an interpolation result of the target hyperparameter;
[0044] A hyperparameter to be verified generation unit, configured to generate a hyperparameter to be verified based on the historical verification results and the interpolation result of the target hyperparameter;
[0045] A hyperparameter sending unit, configured to send the hyperparameter to be verified to a queue;
[0046] A verification result acquisition unit, configured to acquire the verification result of the hyperparameter to be verified from the queue; the verification result of the hyperparameter to be verified is obtained by multiple verification nodes in a distributed cluster acquiring the hyperparameter to be verified from the queue and verifying the acquired hyperparameter to be verified.
[0047] In an exemplary embodiment, the current historical data includes the verification results of currently verified hyperparameters;
[0048] The result interpolation unit includes:
[0049] A verification result calculation unit, configured to perform calculation on the historical verification results corresponding to the historical hyperparameters to obtain a verification result calculation value;
[0050] An interpolation result determination unit, configured to perform determining the verification result calculation value as the interpolation result of the target hyperparameter.
[0051] In an exemplary embodiment, the verification result of the currently verified hyperparameter includes a full - resource verification result obtained by verifying the currently verified hyperparameter based on all resources, and a partial - resource verification result obtained by verifying the currently verified hyperparameter based on partial resources;
[0052] The verification result calculation unit includes:
[0053] A resource quantity determination unit, configured to perform determining the resource quantity of the target hyperparameter; the resource quantity is used to characterize the resource quantity information based on which a hyperparameter is verified;
[0054] A first calculation unit, configured to perform calculating the full - resource verification result to obtain a first calculation value when the resource quantity is all resources;
[0055] A second calculation unit, configured to perform calculating the partial - resource verification result to obtain a second calculation value when the resource quantity is partial resources;
[0056] A first determination unit, configured to perform determining the first calculation value or the second calculation value as the verification result calculation value.
[0057] In an exemplary embodiment, the to - be - verified hyperparameter generation unit includes:
[0058] A first update unit, configured to perform updating the current integrated probability surrogate model based on the historical verification results and the interpolation result of the target hyperparameter to obtain an updated integrated probability surrogate model;
[0059] A first generation unit, configured to perform generating the to - be - verified hyperparameter based on the updated integrated probability surrogate model.
[0060] In an exemplary embodiment, the device includes:
[0061] An initial model acquisition unit, configured to perform acquiring an initial all - resource surrogate model and at least one initial partial - resource surrogate model;
[0062] A weight allocation unit, configured to execute weight allocation for the initial full - scale resource proxy model and the at least one initial partial resource proxy model; wherein the sum of the weight corresponding to the initial full - scale resource proxy model and the weights corresponding to the at least one initial partial resource proxy model is 1;
[0063] A second generation unit, configured to execute generating the integrated probability proxy model based on the initial full - scale resource proxy model, the weight of the initial full - scale resource proxy model, the at least one initial partial resource proxy model, and the weights of the at least one initial partial resource proxy model.
[0064] In an exemplary embodiment, the first update unit includes:
[0065] A second update unit, configured to execute when the imputation result of the target hyperparameter includes the full - scale resource imputation result corresponding to the full - scale resource, updating the current full - scale resource proxy model in the current integrated probability proxy model with the full - scale resource imputation result;
[0066] A third update unit, configured to execute when the imputation result of the target hyperparameter includes the partial resource imputation result corresponding to the partial resource, updating the current partial resource proxy model in the current integrated probability proxy model with the partial resource imputation result;
[0067] A third generation unit, configured to execute obtaining the updated integrated probability proxy model based on the updated current full - scale resource proxy model and the updated current partial resource proxy model.
[0068] In an exemplary embodiment, the apparatus further includes:
[0069] A selection unit, configured to execute selecting at least one set of hyperparameter combinations; the hyperparameter combinations include at least two hyperparameters;
[0070] A second determination unit, configured to execute determining a first verification result corresponding to the at least one set of hyperparameter combinations; the first verification result is a result obtained by verifying the at least one set of hyperparameter combinations based on the full - scale resource;
[0071] A third determination unit, configured to execute determining a second verification result corresponding to the at least one set of hyperparameter combinations; the second verification result is a result obtained by verifying the at least one set of hyperparameter combinations based on the partial resource;
[0072] A consistency check unit, configured to execute consistency checking on the partial order relationship of the first verification result and the partial order relationship of the second verification result;
[0073] A weight adjustment unit, configured to perform weight adjustment of the full - scale resource proxy model and the at least one partial resource proxy model based on the consistency check result.
[0074] In an exemplary embodiment, the apparatus further includes:
[0075] A first resource amount determination unit, configured to perform obtaining a first resource amount based on the weight of the current full - scale resource proxy model and the resource amount corresponding to the current full - scale resource proxy model;
[0076] A second resource amount determination unit, configured to perform obtaining a second resource amount based on the weight of the current partial resource proxy model and the resource amount corresponding to the current partial resource proxy model;
[0077] An initial resource amount determination unit, configured to perform calculating an initial resource amount used for verifying the hyperparameter to be verified according to the first resource amount and the second resource amount.
[0078] According to a third aspect of the embodiments of the present disclosure, there is provided an electronic device, including: a processor; a memory for storing instructions executable by the processor; wherein, the processor is configured to execute the instructions to implement the hyperparameter processing method as described above.
[0079] According to a fourth aspect of the embodiments of the present disclosure, there is provided a computer - readable storage medium, when the instructions in the computer - readable storage medium are executed by a processor of a server, enabling the server to execute the hyperparameter processing method as described above.
[0080] According to a fifth aspect of the embodiments of the present disclosure, there is provided a computer program product, the computer program product includes a computer program, the computer program is stored in a readable storage medium, and at least one processor of a computer device reads and executes the computer program from the readable storage medium, enabling the device to execute the above - mentioned hyperparameter processing method.
[0081] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects:
[0082] The present disclosure obtains historical verification results corresponding to historical hyperparameters, performs verification result interpolation on target hyperparameters for which verification results have not been returned based on the historical verification results, and thus can generate hyperparameters to be verified based on the historical verification results and the interpolation results of the target hyperparameters; sends the hyperparameters to be verified to a queue, and obtains verification results for the hyperparameters to be verified from the queue; here, the verification results of the hyperparameters to be verified are obtained by multiple verification nodes in a distributed cluster respectively obtaining the hyperparameters to be verified from the queue and then verifying the obtained hyperparameters. When there are target hyperparameters for which verification results have not been returned currently, in order to avoid generating the same or similar hyperparameters to be verified based on unupdated historical data, the present disclosure performs interpolation on the verification results of the target hyperparameters, generates hyperparameters to be verified based on the interpolated data, thereby improving the exploration of the hyperparameter space, avoiding waste of verification resources caused by verifying the same or similar hyperparameters to be verified, and improving the efficiency of hyperparameter verification and resource utilization; further, the interpolation results of the target hyperparameters in the present disclosure are obtained based on historical data, so that the interpolation results of the target hyperparameters are adapted to the historical data, can reduce the deviation between the interpolation results of the target hyperparameters and the actual verification results, and make the hyperparameters to be verified generated based on the interpolation results have high verification value.
[0083] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0084] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure, and do not constitute an improper limitation of the present disclosure.
[0085] Figure 1 is a schematic diagram of an implementation environment shown according to an exemplary embodiment.
[0086] Figure 2 is a flowchart of a hyperparameter processing method shown according to an exemplary embodiment.
[0087] Figure 3 is a flowchart of an interpolation result generation method shown according to an exemplary embodiment.
[0088] Figure 4 is a flowchart of a verification result calculated value generation method shown according to an exemplary embodiment.
[0089] Figure 5 is a flowchart of a hyperparameter to be verified generation method shown according to an exemplary embodiment.
[0090] Figure 6It is a flowchart of a method for initializing an integrated probability surrogate model shown according to an exemplary embodiment.
[0091] Figure 7 It is a flowchart of a method for updating an integrated probability surrogate model shown according to an exemplary embodiment.
[0092] Figure 8 It is a flowchart of a method for adjusting model weights shown according to an exemplary embodiment.
[0093] Figure 9 It is a flowchart of a method for determining an initial resource amount shown according to an exemplary embodiment.
[0094] Figure 10 It is a block diagram of a hyperparameter processing device shown according to an exemplary embodiment.
[0095] Figure 11 It is a schematic diagram of a device structure shown according to an exemplary embodiment. Detailed implementation manners
[0096] To enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0097] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0098] Please refer to Figure 1 , which shows a schematic diagram of an implementation environment provided by an embodiment of the present disclosure. The present disclosure can be applied to a distributed cluster, and the distributed cluster can specifically include an optimization node 110 and a plurality of verification nodes 120.
[0099] Specifically, the optimization node 110 generates hyperparameters to be verified according to the corresponding optimization algorithm and sends them to the message queue; the message queue is responsible for caching the hyperparameters waiting to be verified; each verification node 120 can obtain the hyperparameters to be verified from the message queue and verify the corresponding obtained hyperparameters, and send the verification to the message queue; that is, the message queue is also used to cache the verification results generated by the verification node 120; the optimization node 110 obtains the verification results of the verified hyperparameters from the message queue. Among them, the optimization node 110 and the verification node 120 can be physical servers or cloud servers.
[0100] Based on Figure 1 the schematic diagram of the implementation environment, the distributed hyperparameter optimization framework based on the message queue in the present disclosure can be determined. This framework mainly includes three parts:
[0101] 1. Optimization node (optimizer): responsible for giving the hyperparameters that need to be verified;
[0102] 2. Message queue: responsible for caching the hyperparameters waiting to be verified and the verification results waiting to collect the hyperparameters; specifically, it can include a waiting queue and a result queue;
[0103] 3. Verification node (verifier, worker node): verifies the hyperparameters to be verified obtained from the message queue to obtain the verification results.
[0104] The running process of the entire distributed framework includes:
[0105] 1. The optimization node generates hyperparameter recommendations and sends them to the waiting queue; the queue manager records the status of these hyperparameters as waiting to be verified.
[0106] 2. The verification node obtains the hyperparameters from the waiting queue for verification, and the queue manager records the status of the hyperparameters obtained by the verified node as being verified.
[0107] 3. The verification node sends the verification results to the result queue, and the queue manager records the status of the verified hyperparameters as verified.
[0108] 4. The optimization node obtains the verification results from the result queue, updates the current historical data, and makes the next round of hyperparameter recommendations based on the updated historical data.
[0109] Due to the use of the waiting queue and the result queue for buffering, the optimization node does not need to know the specific information of the verification nodes. It only needs to interact with the two queues, so the number of verification nodes can be freely increased or decreased. In addition, based on this decoupling characteristic, the distributed framework has good fault tolerance. If a certain verification node fails when verifying the hyperparameter a, since its result is not fed back to the result queue within the specified time, the hyperparameter a will be reset to waiting for verification in the waiting queue again and wait for the verification node to obtain and verify it later. If the optimization node or the queue fails, it can be directly restored through the persisted and backed-up verification history. Through the abstract structure of the message queue, the hyperparameter optimization task is decoupled and assigned to different verification nodes in the distributed cluster, improving resource utilization.
[0110] It should be noted that the optimization algorithm in the present disclosure can be either a synchronous parallel algorithm or an asynchronous parallel algorithm. For the synchronous parallel optimization algorithm, the optimizer gives multiple hyperparameters (referred to as a set of hyperparameters) in each round and waits for all the verification results of this set of hyperparameters to be returned before giving the recommendation for the next set of hyperparameters.
[0111] For the asynchronous parallel optimization algorithm, whenever the verifier is in the working state, the optimizer gives the recommendation for the next hyperparameter configuration. Among them, each hyperparameter configuration can include parameters of at least one dimension. For example, a hyperparameter configuration can include parameters of multiple dimensions such as the model learning rate, the type of activation function, and the number of hidden layers of the deep neural network; in a specific embodiment, there can be a hyperparameter configuration (a, b, c), where a corresponds to the model learning rate, b corresponds to the type of activation function of the model, and c corresponds to the number of hidden layers of the deep neural network. By combining different values of a, b, and c, different hyperparameter configurations will be obtained.
[0112] The optimizer can judge whether the verifier is in the working state according to the result returned by the verifier. Or for the situation where the verifier goes online and offline dynamically, the verifier can inform the optimizer whether it will work next in the result. Here, the working state specifically refers to that the verifier has currently returned the verification result and whether it will perform the hyperparameter verification work next. If so, it is judged that it is in the working state, that is, the online state; if not, for example, the verifier goes offline after returning the current verification result, then it is not in the working state.
[0113] To solve the problems of the exploratory decline of the hyperparameter space and the waste of verification resources in the related art, the embodiments of the present disclosure provide a hyperparameter processing method, and its execution subject can be Figure 1 the optimization node in Figure 2 , specifically, please refer to
[0114] S210. Obtain the historical verification results corresponding to the historical hyperparameters; the historical verification results are used to characterize the performance data of the trained model corresponding to the historical hyperparameters; the trained model corresponding to the historical hyperparameters is obtained by training the model to be trained based on the historical hyperparameters.
[0115] The historical verification results are used to characterize the performance data of the trained model corresponding to the historical hyperparameters. The performance data of the trained model can be represented by relevant evaluation metrics, and the evaluation metrics can be the accuracy, recall rate, true positive rate, false positive rate, etc. of the model.
[0116] In a specific embodiment, when the model to be trained is an image classification model, the historical verification results corresponding to the historical hyperparameters can be: the classification accuracy of the model obtained after training the image classification model based on the historical hyperparameters.
[0117] In an alternative embodiment, when the model to be trained is an object recognition model, the historical verification results corresponding to the historical hyperparameters can be: the recognition accuracy of the model obtained after training the object recognition model based on the historical hyperparameters.
[0118] S220. Perform verification result interpolation on the target hyperparameters that have not returned verification results based on the historical verification results to obtain the interpolation results of the target hyperparameters.
[0119] Data interpolation refers to the processing method of adding data in the case of missing data to obtain a complete data set; in this embodiment, since the target hyperparameters are in the verification state and no verification results are returned, and the next generation of hyperparameters to be verified requires the verification results of the target hyperparameters, the method of performing verification result interpolation on the target hyperparameters is thus adopted to fill the data. When interpolating the verification results of the target hyperparameters, it is necessary to perform interpolation based on the current existing historical verification result data distribution, so that the interpolated data will not be too obtrusive and conform to the distribution of historical data.
[0120] In the process of parallel hyperparameter optimization and verification, whether it is a synchronous or asynchronous optimization algorithm, when new hyperparameter recommendations need to be given, there may be hyperparameters that are being verified but have not returned results, or hyperparameters that are cached in the waiting queue waiting for verification, and neither of these two types of hyperparameters has returned verification results. For the synchronous optimization algorithm, in the process of the optimizer giving a set of hyperparameter recommendations, if this set of hyperparameters is viewed from the perspective of sequence, when giving subsequent hyperparameter recommendations, the hyperparameters recommended earlier are those that have not returned results. For the asynchronous optimization algorithm, since the verifier is idle whenever a result is returned and can obtain new hyperparameters from the waiting queue, at this time, there may be other verifiers that are verifying hyperparameters but have not returned results.
[0121] The existence of hyperparameters for which no results are returned may lead to a decline in the parallel recommendation effect of the optimization algorithm. This is because when the optimization algorithm recommends hyperparameters based on current historical data, specifically the historical verification results, if multiple sets of hyperparameters are required while the historical verification results remain unchanged, these hyperparameters may be the same or very similar, resulting in a decrease in the exploration of the hyperparameter space by the optimization algorithm and a waste of verification resources. Therefore, in the present disclosure, a verification result interpolation method is introduced, specifically, interpolation is performed on the target hyperparameters for which no verification results are returned to obtain interpolation results.
[0122] The target hyperparameters in the embodiments of the present disclosure may be one or more. Hereinafter, the case where the target hyperparameter is one will be described. When there are multiple target hyperparameters, the corresponding implementation methods are similar.
[0123] Please refer to Figure 3 , which shows a method for generating interpolation results. The method may include:
[0124] S310. Calculate the verification result calculation value for the historical verification results corresponding to the historical hyperparameters.
[0125] S320. Determine the verification result calculation value as the interpolation result of the target hyperparameter.
[0126] The historical hyperparameters may be verified hyperparameters, and one verified hyperparameter corresponds to one verification result. When calculating the interpolation result of the target hyperparameter, it is necessary to perform calculations based on the verification results of the verified hyperparameters. For example, statistical calculations may be performed on the verification results of each verified hyperparameter. Specifically, the median of the verification results of each verified hyperparameter may be calculated, or the average value of the verification results of each verified hyperparameter may be calculated, etc. In the present disclosure, the interpolation result of the target hyperparameter is obtained based on historical data, making the interpolation result of the target hyperparameter adaptable to the historical data, reducing the deviation between the interpolation result of the target hyperparameter and the actual verification result, and making the hyperparameters to be verified generated based on the interpolation result have a high verification value; and by performing corresponding statistical value calculations on the verification results of the currently verified hyperparameters, the calculation process is simple, thereby being able to improve the verification result interpolation efficiency accordingly.
[0127] Further, the verification results of the currently verified hyperparameters include the full - resource verification results obtained by verifying the currently verified hyperparameters based on all resources, and the partial - resource verification results obtained by verifying the currently verified hyperparameters based on partial resources. The resources in the embodiments of the present disclosure may be the amount of resource data required for verification, the number of iteration rounds during the verification process, etc., and may be specifically set according to the specific implementation situation.
[0128] In a specific embodiment, if the resource is specifically the amount of resource data required for verification, and the amount of resource data is M, then the corresponding full amount of resource is the amount of resource data M. The partial resource can be 1 / 3 resource, 1 / 9 resource, etc. Correspondingly, the 1 / 3 resource is the amount of resource data M / 3, and the 1 / 9 resource is the amount of resource data M / 9.
[0129] In a specific embodiment, if the resource is specifically the number of iteration rounds in the verification process, and the number of iteration rounds is N, then the corresponding full amount of resource is N, the 1 / 3 resource is the amount of resource data M / 3, and the 1 / 9 resource is the amount of resource data M / 9.
[0130] Please refer to Figure 4 , which shows a method for generating a verification result calculated value. The method may include:
[0131] S410. Determine the amount of resources for the target hyperparameter; the amount of resources is used to represent the resource quantity information based on which a hyperparameter is verified.
[0132] The amount of resources here can be the above-mentioned amount of resource data or the number of iteration rounds, etc.
[0133] S420. When the amount of resources is the full amount of resources, calculate the verification result for the full amount of resources to obtain a first calculated value.
[0134] S430. When the amount of resources is partial resources, calculate the verification result for the partial resources to obtain a second calculated value.
[0135] S440. Determine the verification result calculated value as the first calculated value or the second calculated value.
[0136] In an alternative embodiment, since there may be verification results obtained based on different amounts of resources, when performing result interpolation for the target hyperparameter, it is necessary to first determine the verification result corresponding to the corresponding amount of resources. For example, when the amount of resources used for verifying the target hyperparameter without a returned verification result is 1 / 3, there will be a verification result corresponding to the 1 / 3 amount of resources. Based on this verification result, interpolation is performed for the target hyperparameter, further improving the accuracy and adaptability of the result interpolation.
[0137] S230. Generate a hyperparameter to be verified based on the historical verification result and the interpolation result of the target hyperparameter.
[0138] Please refer to Figure 5 , which shows a method for generating a hyperparameter to be verified. The method may include:
[0139] S510. Update the current integrated probability surrogate model based on the historical verification results and the imputation results of the target hyperparameters to obtain an updated integrated probability surrogate model.
[0140] S520. Generate the hyperparameters to be verified based on the updated integrated probability surrogate model.
[0141] In a specific embodiment, the Bayesian optimization algorithm is taken as an example for illustration. The Bayesian optimization algorithm is an iterative optimization framework based on a model, which includes two important components, namely, a probabilistic surrogate model and an acquisition function. The main steps of the optimization are as follows:
[0142] 1. Based on the existing historical observation data, use the probabilistic surrogate model to model the input and output of the hyperparameter configuration.
[0143] 2. The surrogate model predicts the probability distribution of candidate points in the input space, and the acquisition function calculates the verification value of the candidate points according to the probabilistic surrogate model.
[0144] 3. Optimize the acquisition function to obtain the next candidate point with the highest value, verify the parameter configuration using the verification function, and update the result to the historical observation data.
[0145] 4. Repeat the above steps 1-3 until the given resource constraint is reached or the expected effect is achieved.
[0146] Take the target hyperparameters as the input of the current integrated probability surrogate model, and take the imputation results of the target hyperparameters as the output of the current integrated probability surrogate model to update the current integrated probability surrogate model; then use the acquisition function to determine the next hyperparameters to be verified based on the predicted point probability distribution shown by the updated current integrated probability surrogate model.
[0147] In the present disclosure, the integrated probability surrogate model includes surrogate models corresponding to multiple different resource amounts. For example, a surrogate model corresponding to the full amount of resources, a surrogate model corresponding to partial resources. The surrogate model corresponding to partial resources can specifically be a surrogate model corresponding to 1 / 3 of the resources, a surrogate model corresponding to 1 / 9 of the resources, etc.
[0148] Please refer to Figure 6 , which shows a method for initializing an integrated probability surrogate model. The method may include:
[0149] S610. Obtain an initial full-resource surrogate model and at least one initial partial-resource surrogate model.
[0150] S620. Assign weights to the initial full - scale resource proxy model and the at least one initial partial - scale resource proxy model; where the sum of the weights corresponding to the initial full - scale resource proxy model and the weights corresponding to the at least one initial partial - scale resource proxy model is 1.
[0151] S630. Generate the integrated probability proxy model based on the initial full - scale resource proxy model, the weight of the initial full - scale resource proxy model, the at least one initial partial - scale resource proxy model, and the weights of the at least one initial partial - scale resource proxy model.
[0152] In a specific embodiment, both the initial full - scale resource proxy model and the initial partial - scale resource proxy model can adopt Gaussian models, and the initial weights of the corresponding models can be assigned based on empirical values.
[0153] In the integrated probability proxy model, a separate proxy model is trained for the information of each precision. Here, the precision can specifically refer to the resource quantity. Each proxy model calculates the weight based on reliability. When performing hyperparameter recommendation, the influence of each sub - model in the integrated probability proxy model can be determined according to the weight of each proxy model. In the initial stage of optimization operation, there is little historical information for complete verification. By using information of multiple precisions, the optimization algorithm can better learn the relationship between hyperparameters and verification results, so as to give better hyperparameter recommendations and accelerate hyperparameter optimization.
[0154] The integrated probability proxy model can be updated accordingly based on the verification results of the verified hyperparameters. Among them, each sub - model corresponding to each resource quantity in the integrated probability proxy model needs to be updated accordingly based on the verification results of the corresponding resource quantity; please refer to Figure 7 , which shows a method for updating the integrated probability proxy model. The method may include:
[0155] S710. When the interpolation result of the target hyperparameter includes the full - scale resource interpolation result corresponding to the full - scale resource, use the full - scale resource interpolation result to update the current full - scale resource proxy model in the current integrated probability proxy model.
[0156] S720. When the interpolation result of the target hyperparameter includes the partial - scale resource interpolation result corresponding to the partial - scale resource, use the partial - scale resource interpolation result to update the current partial - scale resource proxy model in the current integrated probability proxy model.
[0157] S730. Obtain the updated integrated probability proxy model based on the updated current full - scale resource proxy model and the updated current partial - scale resource proxy model.
[0158] In a specific embodiment, when updating the full - scale resource proxy model, it is necessary to use the full - scale resource interpolation result obtained by interpolation according to the full - scale resource verification result for the update; when updating the partial resource proxy model, it is necessary to use the partial resource interpolation result obtained by interpolation according to the partial resource verification result for the update.
[0159] In a specific embodiment, if the verification result of the verified hyperparameters is currently returned, and the verification result of the verified hyperparameters is obtained by verifying based on the full - scale resources, then use the verification result of the verified hyperparameters to update the full - scale resource proxy model; if the verification result of the verified hyperparameters is obtained by verifying based on the partial resources, then use the verification result of the verified hyperparameters to update the partial resource proxy model.
[0160] Furthermore, the weights corresponding to each sub - model in the integrated probability proxy model can be adjusted. Please refer to Figure 8 , which shows a model weight adjustment method. This method may include:
[0161] S810. Select at least one set of hyperparameter combinations; the hyperparameter combinations include at least two hyperparameters.
[0162] S820. Determine the first verification result corresponding to the at least one set of hyperparameter combinations; the first verification result is the result obtained by verifying the at least one set of hyperparameter combinations based on the full - scale resources.
[0163] S830. Determine the second verification result corresponding to the at least one set of hyperparameter combinations; the second verification result is the result obtained by verifying the at least one set of hyperparameter combinations based on the partial resources.
[0164] S840. Perform a consistency check on the partial order relationship of the first verification result and the partial order relationship of the second verification result.
[0165] S850. Adjust the weights of the full - scale resource proxy model and the at least one partial resource proxy model based on the consistency check result.
[0166] In a specific embodiment, the historical data corresponding to the full amount of resources includes L verified hyperparameters and the corresponding L hyperparameter verification results. Pairwise combination of the L hyperparameters can obtain L(L - 1) / 2 combinations. For example, for the hyperparameter combination [a, b], the corresponding verification results are [ya1, yb1]; find hyperparameters a and b in the historical data corresponding to the partial resources, determine the corresponding verification results as [ya2, yb2], and judge whether the partial order relationship between [ya1, yb1] and [ya2, yb2] is consistent. The partial order relationship here can be a magnitude relationship. For example, it can be understood that if ya1 > yb1, judge whether ya2 is greater than yb2. If so, judge that the partial order relationship between the two is consistent; if not, judge that the partial order relationship between the two is inconsistent. For the L(L - 1) / 2 combinations, the above partial order relationship consistency check can be performed, and the corresponding model weights can be determined according to the number of hyperparameter combinations passing the consistency check. Specifically, when adjusting the weights, it can be determined according to the number of combinations passing the consistency check. If it is greater than the preset value, the weight of the partial resource proxy model can be increased; if not, the weight of the partial resource proxy model can be correspondingly decreased. Thus, the weights of each sub-model are adaptively adjusted according to the verification results, so that the generated integrated model can be more applicable to the current hyperparameter verification scenario and improve the efficiency of hyperparameter verification.
[0167] In the multi-precision optimization based on early stopping, since the confidence level of the verification results corresponding to a smaller amount of resources may be low and cannot fully reflect the verification results using the complete amount of resources, that is, the hyperparameter configuration with worse verification performance using a smaller amount of resources may have better verification performance using the complete amount of resources. If the initial amount of resources is fixed and the confidence level of the verification results of this amount of resources is low, the hyperparameter configuration with better early stopping effect may be wrongly stopped, resulting in a decline in the optimization effect. Please refer to Figure 9 , which shows a method for determining the initial amount of resources. The method may include:
[0168] S910. Obtain a first amount of resources based on the weight of the current full-resource proxy model and the amount of resources corresponding to the current full-resource proxy model.
[0169] S920. Obtain a second amount of resources based on the weight of the current partial-resource proxy model and the amount of resources corresponding to the current partial-resource proxy model.
[0170] S930. Calculate the initial amount of resources used for verifying the hyperparameters to be verified according to the first amount of resources and the second amount of resources.
[0171] In a specific embodiment, the integrated probability proxy model includes a full - scale resource proxy model and a 1 / 3 - scale resource proxy model. The weight corresponding to the full - scale resource proxy model is 0.7, and the weight corresponding to the 1 / 3 - scale resource proxy model is 0.3. Accordingly, the initial resource amount can be determined by means of probability sampling. Specifically, after probability sampling, there is a 70% probability that the initial resource amount is the full - scale resource, and a 30% probability that the initial resource amount is the 1 / 3 - scale resource amount.
[0172] Each sub - model with a certain accuracy (i.e., each verified resource amount) in the integrated probability proxy model has a corresponding weight, which reflects the confidence level of the verification result of using the corresponding resource amount to verify the hyper - parameters. Therefore, by using the weights of each sub - model, the initial resource amount of the new hyper - parameter configuration samples the model corresponding resource amount according to the model weight ratio, reducing the time waste caused by the configuration verification under the low - confidence resource amount and further improving the efficiency of the multi - accuracy optimization algorithm.
[0173] S240. Send the hyper - parameters to be verified to the queue.
[0174] S250. Obtain the verification result of the hyper - parameters to be verified from the queue; the verification result of the hyper - parameters to be verified is obtained by multiple verification nodes in the distributed cluster obtaining the hyper - parameters to be verified from the queue and verifying the obtained hyper - parameters to be verified.
[0175] In an alternative embodiment, in algorithms such as Bayesian optimization, it is required to train each hyper - parameter based on the full - scale resources (for example, using the complete data set, or training the deep model for the complete number of iterations), and give the performance result of the model on the validation set. After obtaining the hyper - parameters to be verified, the verification node defaults to perform a complete training and returns the performance result of the trained model, that is, the complete verification result.
[0176] In an alternative embodiment, for a given hyper - parameter, it is required that the model be trained on the given resource amount and then give the performance result of the model on the validation set. Since the given resource amount may be less than the complete resource amount (for example, training the model using a part of the data set, or only training the deep - learning model for a small number of iterations), the returned result is called a "partial verification result", and the verification process is called "partial verification".
[0177] In a distributed parallel hyperparameter optimization framework, for a multi-precision optimization algorithm, while giving hyperparameter recommendations, the optimizer also provides the amount of resources required for verifying this hyperparameter configuration. After obtaining the hyperparameter configuration and the corresponding amount of resources, the verification node can perform "partial verification" on the hyperparameters according to the amount of resources. For example, the dataset is sampled according to the proportion of the amount of resources and then used to train and verify the model, or the deep learning model is trained for a certain number of iterations according to the amount of resources and then the current model performance is returned.
[0178] In addition, according to the multi-precision algorithm, for the same hyperparameter, the optimizer may require the verifier to perform verifications on different amounts of resources successively. For a deep learning model, when the verifier supports training the model for more iterations, the model with the same hyperparameter that was previously trained for fewer iterations can continue to be trained, thereby reducing the overhead caused by retraining.
[0179] In an optional embodiment, the optimizer can use the following optimization algorithms:
[0180] 1. By introducing a median imputation strategy, the Bayesian optimization algorithm (serial algorithm) is extended to run in parallel, thereby supporting (synchronous, asynchronous) parallel Bayesian optimization algorithms.
[0181] 2. Multi-precision synchronous parallel optimization methods based on early stopping mechanisms such as HyperBand and BOHB.
[0182] 3. Asynchronous parallel multi-precision optimization methods based on boosting mechanisms.
[0183] 4. Synchronous and asynchronous parallel multi-precision optimization methods based on multi-precision Bayesian optimization.
[0184] For the method in (4), the initial resource method of sampling according to the weights of the multi-precision Bayesian model is further used to accelerate the parallel efficiency of multi-precision optimization. If the optimization algorithm is a multi-precision optimization algorithm, in the configuration sent by the optimizer, in addition to the hyperparameter configuration, the resource amount parameter required for verification also needs to be included. If the optimization algorithm is not a multi-precision optimization algorithm, the optimizer only needs to send the hyperparameter configuration.
[0185] In an optional embodiment, for hyperparameters whose results have not been returned yet, the introduced verification result number imputation algorithm sets their verification results to the median of the existing verification results and adds them to the verified historical data for the Bayesian optimization algorithm to perform parallel hyperparameter recommendations; this method improves the speed of the Bayesian optimization algorithm through parallelization while ensuring that the convergence is not reduced, and can find better hyperparameters in a shorter time compared to the serial Bayesian optimization algorithm.
[0186] In an alternative embodiment, algorithms such as Hyperband and BOHB are essentially synchronous parallel algorithms, which suffer from waiting and waste of machine resources. The optimizer also supports an asynchronous parallel multi-precision optimization method based on a boosting mechanism. In this method, the early stopping judgment no longer waits for all hyperparameter configuration results in a round to be returned before proceeding. Instead, at any moment when a validator is idle, the hyperparameter configurations that rank in the top 1 / eta (eta is usually set to 2, 3, or 4) in terms of performance among the hyperparameter configurations being verified with the same amount of resources are "boosted", that is, more resources are given to the validator for verification. Through this method, the problem of resource waste caused by the waiting of validator nodes is avoided, and distributed parallel resources are more fully utilized to accelerate hyperparameter optimization.
[0187] The present disclosure proposes an abstract structure of a message queue, decouples hyperparameter optimization tasks, and allocates them to different validation nodes, enabling parallel verification of hyperparameters using distributed resources and improving resource utilization; supports synchronous and asynchronous parallel operation of various hyperparameter optimization algorithms, including Bayesian optimization algorithms, multi-precision optimization algorithms, etc., to improve the search efficiency of hyperparameter optimization (increase the optimization speed); introduces a median imputation strategy in hyperparameter search and supports both synchronous and asynchronous parallel modes; uses a method of sampling initial resources according to the weights of a multi-precision Bayesian model to accelerate the parallel efficiency of multi-precision optimization; thus, through the solution of the present disclosure, better hyperparameter configurations can be searched in a shorter time, achieving better hyperparameter optimization effects.
[0188] Figure 10 is a block diagram of a hyperparameter processing device shown according to an exemplary embodiment. Referring to Figure 10 , the device includes:
[0189] A historical verification result acquisition unit 1010, configured to acquire historical verification results corresponding to historical hyperparameters; the historical verification results are used to characterize the performance data of the trained model corresponding to the historical hyperparameters; the trained model corresponding to the historical hyperparameters is obtained by training a model to be trained based on the historical hyperparameters;
[0190] A result imputation unit 1020, configured to perform verification result imputation on target hyperparameters for which verification results have not been returned based on the historical verification results to obtain imputation results of the target hyperparameters;
[0191] A hyperparameter to be verified generation unit 1030, configured to generate hyperparameters to be verified based on the historical verification results and the imputation results of the target hyperparameters;
[0192] A hyperparameter sending unit 1040, configured to send the hyperparameters to be verified to a queue;
[0193] The verification result acquisition unit 1050 is configured to obtain the verification result of the hyperparameter to be verified from the queue; the verification result of the hyperparameter to be verified is obtained by multiple verification nodes in the distributed cluster from the queue to obtain the hyperparameter to be verified and verify the obtained hyperparameter to be verified.
[0194] In an exemplary embodiment,
[0195] The result interpolation unit 1010 includes:
[0196] The verification result calculation unit is configured to calculate the historical verification result corresponding to the historical hyperparameter to obtain a verification result calculation value;
[0197] The interpolation result determination unit is configured to determine the verification result calculation value as the interpolation result of the target hyperparameter.
[0198] In an exemplary embodiment, the verification result of the currently verified hyperparameter includes a full - resource verification result obtained by verifying the currently verified hyperparameter based on all resources, and a partial - resource verification result obtained by verifying the currently verified hyperparameter based on partial resources;
[0199] The verification result calculation unit includes:
[0200] The resource amount determination unit is configured to determine the resource amount of the target hyperparameter; the resource amount is used to represent the resource quantity information based on which a hyperparameter is verified.
[0201] The first calculation unit is configured to calculate the full - resource verification result to obtain a first calculation value when the resource amount is all resources;
[0202] The second calculation unit is configured to calculate the partial - resource verification result to obtain a second calculation value when the resource amount is partial resources;
[0203] The first determination unit is configured to determine the first calculation value or the second calculation value as the verification result calculation value.
[0204] In an exemplary embodiment, the hyperparameter to be verified generation unit 1020 includes:
[0205] The first update unit is configured to update the current integrated probability proxy model based on the historical verification result and the interpolation result of the target hyperparameter to obtain an updated integrated probability proxy model;
[0206] The first generation unit is configured to generate the hyperparameters to be verified based on the updated integrated probability surrogate model.
[0207] In an exemplary embodiment, the apparatus includes:
[0208] An initial model acquisition unit configured to acquire an initial full - scale resource surrogate model and at least one initial partial - scale resource surrogate model;
[0209] A weight assignment unit configured to assign weights to the initial full - scale resource surrogate model and the at least one initial partial - scale resource surrogate model; wherein the sum of the weights corresponding to the initial full - scale resource surrogate model and the weights corresponding to the at least one initial partial - scale resource surrogate model is 1;
[0210] A second generation unit configured to generate the integrated probability surrogate model based on the initial full - scale resource surrogate model, the weight of the initial full - scale resource surrogate model, the at least one initial partial - scale resource surrogate model, and the weights of the at least one initial partial - scale resource surrogate model.
[0211] In an exemplary embodiment, the first update unit includes:
[0212] A second update unit configured to, when the interpolation result of the target hyperparameter includes a full - scale resource interpolation result corresponding to the full - scale resource, update the current full - scale resource surrogate model in the current integrated probability surrogate model with the full - scale resource interpolation result;
[0213] A third update unit configured to, when the interpolation result of the target hyperparameter includes a partial - scale resource interpolation result corresponding to the partial - scale resource, update the current partial - scale resource surrogate model in the current integrated probability surrogate model with the partial - scale resource interpolation result;
[0214] A third generation unit configured to obtain the updated integrated probability surrogate model based on the updated current full - scale resource surrogate model and the updated current partial - scale resource surrogate model.
[0215] In an exemplary embodiment, the apparatus further includes:
[0216] A selection unit configured to select at least one set of hyperparameter combinations; the hyperparameter combinations include at least two hyperparameters;
[0217] A second determination unit configured to determine a first verification result corresponding to the at least one set of hyperparameter combinations; the first verification result is a result obtained by verifying the at least one set of hyperparameter combinations based on the full - scale resources.
[0218] A third determination unit, configured to perform a determination of a second verification result corresponding to the at least one set of hyperparameter combinations; the second verification result is a result obtained by verifying the at least one set of hyperparameter combinations based on partial resources.
[0219] A consistency check unit, configured to perform a consistency check on the partial order relationship of the first verification result and the partial order relationship of the second verification result.
[0220] A weight adjustment unit, configured to perform an adjustment of the weights of the full - scale resource proxy model and the at least one partial - resource proxy model based on the consistency check result.
[0221] In an exemplary embodiment, the apparatus further includes:
[0222] A first resource amount determination unit, configured to perform an operation of obtaining a first resource amount based on the weight of the current full - scale resource proxy model and the resource amount corresponding to the current full - scale resource proxy model.
[0223] A second resource amount determination unit, configured to perform an operation of obtaining a second resource amount based on the weight of the current partial - resource proxy model and the resource amount corresponding to the current partial - resource proxy model.
[0224] An initial resource amount determination unit, configured to perform an operation of calculating an initial resource amount used for verifying the hyperparameters to be verified according to the first resource amount and the second resource amount.
[0225] Regarding the apparatus in the above - mentioned embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.
[0226] In an exemplary embodiment, there is also provided a computer - readable storage medium including instructions. Optionally, the computer - readable storage medium may be a ROM, a random access memory (RAM), a CD - ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.; when the instructions in the computer - readable storage medium are executed by a processor of a server, the server is enabled to execute any of the methods in the above - mentioned embodiments.
[0227] In an exemplary embodiment, there is also provided a computer program product. The computer program product includes a computer program. The computer program is stored in a readable storage medium, and at least one processor of a computer device reads and executes the computer program, enabling the device to execute any of the methods in the above - mentioned embodiments.
[0228] This embodiment also provides a device, and the structure diagram thereof can be referred to Figure 11, the device 1100 can vary significantly due to different configurations or performances. It may include one or more central processing units (CPUs) 1122 (e.g., one or more processors) and a memory 1132, and one or more storage media 1130 (e.g., one or more mass storage devices) for storing application programs 1142 or data 1144. Among them, the memory 1132 and the storage media 1130 can be transient storage or persistent storage. The programs stored in the storage media 1130 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the device. Further, the central processing unit 1122 can be configured to communicate with the storage media 1130 and execute a series of instruction operations in the storage media 1130 on the device 1100. The device 1100 may also include one or more power supplies 1126, one or more wired or wireless network interfaces 1150, one or more input / output interfaces 1158, and / or one or more operating systems 1141, such as Windows Server TM , Mac OS X TM , Unix TM , Linux TM , FreeBSD TM and so on. Any of the above methods in this embodiment can be implemented based on the device shown in Figure 11 .
[0229] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present disclosure. This application is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include common general knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.
[0230] It should be understood that the present disclosure is not limited to the exact structures already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A method for processing hyperparameters, characterized in that, the method includes: Obtaining historical verification results corresponding to historical hyperparameters; the historical verification results are used to characterize the performance data of the trained model corresponding to the historical hyperparameters; the trained model corresponding to the historical hyperparameters is obtained by training the model to be trained based on the historical hyperparameters; when the model to be trained is an image classification model, the historical verification result corresponding to the historical hyperparameters is: the classification accuracy of the model obtained after training the image classification model based on the historical hyperparameters; when the model to be trained is an object recognition model, the historical verification result corresponding to the historical hyperparameters is: the recognition accuracy of the model obtained after training the object recognition model based on the historical hyperparameters; Performing verification result interpolation on the target hyperparameters that have not returned verification results based on the historical verification results to obtain the interpolation results of the target hyperparameters; Generating hyperparameters to be verified based on the historical verification results and the interpolation results of the target hyperparameters; Sending the hyperparameters to be verified to a queue; Obtaining the verification results of the hyperparameters to be verified from the queue; the verification results of the hyperparameters to be verified are obtained by multiple verification nodes in a distributed cluster obtaining the hyperparameters to be verified from the queue and verifying the obtained hyperparameters to be verified.
2. A method for processing hyperparameters according to claim 1, characterized in that, the performing verification result interpolation on the target hyperparameters based on the historical verification results to obtain the interpolation results of the target hyperparameters includes: Calculating the historical verification results corresponding to the historical hyperparameters to obtain a calculated value of the verification result; Determining the calculated value of the verification result as the interpolation result of the target hyperparameters.
3. A method for processing hyperparameters according to claim 2, characterized in that, the historical verification results include full-resource verification results obtained by verifying the historical hyperparameters based on full resources and partial-resource verification results obtained by verifying the historical hyperparameters based on partial resources; the calculating the historical verification results corresponding to the historical hyperparameters to obtain a calculated value of the verification result includes: Determining the resource amount of the target hyperparameters; the resource amount is used to characterize the resource quantity information based on which a hyperparameter is verified; When the resource amount is full resources, calculating the full-resource verification results to obtain a first calculated value; When the resource amount is partial resources, calculating the partial-resource verification results to obtain a second calculated value; Determining the first calculated value or the second calculated value as the calculated value of the verification result.
4. A method for processing hyperparameters according to claim 1, characterized in that, the generating hyperparameters to be verified based on the historical verification results and the interpolation results of the target hyperparameters includes: Updating the current integrated probability proxy model based on the historical verification results and the interpolation results of the target hyperparameters to obtain an updated integrated probability proxy model; Generate the hyperparameters to be verified based on the updated integrated probability surrogate model.
5. A hyperparameter processing method according to claim 4, wherein, the method further includes: obtain an initial full - scale resource surrogate model and at least one initial partial - scale resource surrogate model; assign weights to the initial full - scale resource surrogate model and the at least one initial partial - scale resource surrogate model; wherein the sum of the weights corresponding to the initial full - scale resource surrogate model and the weights corresponding to the at least one initial partial - scale resource surrogate model is 1; generate the integrated probability surrogate model based on the initial full - scale resource surrogate model, the weight of the initial full - scale resource surrogate model, the at least one initial partial - scale resource surrogate model, and the weights of the at least one initial partial - scale resource surrogate model.
6. A hyperparameter processing method according to claim 5, wherein, the updating the current integrated probability surrogate model based on the historical verification result and the imputation result of the target hyperparameter to obtain an updated integrated probability surrogate model includes: when the imputation result of the target hyperparameter includes a full - scale resource imputation result corresponding to full - scale resources, update the current full - scale resource surrogate model in the current integrated probability surrogate model using the full - scale resource imputation result; when the imputation result of the target hyperparameter includes a partial - scale resource imputation result corresponding to partial resources, update the current partial - scale resource surrogate model in the current integrated probability surrogate model using the partial - scale resource imputation result; obtain the updated integrated probability surrogate model based on the updated current full - scale resource surrogate model and the updated current partial - scale resource surrogate model.
7. A hyperparameter processing method according to claim 5, wherein, the method further includes: select at least one set of hyperparameter combinations; the hyperparameter combinations include at least two hyperparameters; determine a first verification result corresponding to the at least one set of hyperparameter combinations; the first verification result is a result obtained by verifying the at least one set of hyperparameter combinations based on full - scale resources; determine a second verification result corresponding to the at least one set of hyperparameter combinations; the second verification result is a result obtained by verifying the at least one set of hyperparameter combinations based on partial resources; perform a consistency check on the partial order relationship of the first verification result and the partial order relationship of the second verification result; adjust the weights of the full - scale resource surrogate model and the at least one partial - scale resource surrogate model based on the consistency check result.
8. A hyperparameter processing method according to claim 6 or 7, wherein, the method further includes: obtain a first resource amount based on the weight of the current full - scale resource surrogate model and the resource amount corresponding to the current full - scale resource surrogate model; obtain a second resource amount based on the weight of the current partial - scale resource surrogate model and the resource amount corresponding to the current partial - scale resource surrogate model; calculate the initial resource amount used for verifying the hyperparameters to be verified according to the first resource amount and the second resource amount.
9. A hyperparameter processing device, wherein, The device includes: A historical verification result acquisition unit, configured to acquire historical verification results corresponding to historical hyperparameters; the historical verification results are used to characterize the performance data of a trained model corresponding to the historical hyperparameters; the trained model corresponding to the historical hyperparameters is obtained by training a model to be trained based on the historical hyperparameters; when the model to be trained is an image classification model, the historical verification result corresponding to the historical hyperparameters is: the classification accuracy of the model obtained after training the image classification model based on the historical hyperparameters; when the model to be trained is an object recognition model, the historical verification result corresponding to the historical hyperparameters is: the recognition accuracy of the model obtained after training the object recognition model based on the historical hyperparameters. A result interpolation unit, configured to perform result interpolation on target hyperparameters for which verification results have not been returned based on the historical verification results to obtain interpolation results of the target hyperparameters. A hyperparameter to be verified generation unit, configured to generate hyperparameters to be verified based on the historical verification results and the interpolation results of the target hyperparameters. A hyperparameter sending unit, configured to send the hyperparameters to be verified to a queue. A verification result acquisition unit, configured to acquire verification results of the hyperparameters to be verified from the queue; the verification results of the hyperparameters to be verified are obtained by multiple verification nodes in a distributed cluster acquiring the hyperparameters to be verified from the queue and verifying the acquired hyperparameters to be verified.
10. An apparatus for processing hyperparameters according to claim 9, wherein, the result interpolation unit includes: A verification result calculation unit, configured to calculate the historical verification results corresponding to the historical hyperparameters to obtain a verification result calculation value. An interpolation result determination unit, configured to determine the verification result calculation value as the interpolation result of the target hyperparameters.
11. An apparatus for processing hyperparameters according to claim 10, wherein, the verification results of the currently verified hyperparameters include full-resource verification results obtained by verifying the currently verified hyperparameters based on all resources and partial-resource verification results obtained by verifying the currently verified hyperparameters based on partial resources. The verification result calculation unit includes: A resource amount determination unit, configured to determine the resource amount of the target hyperparameters; the resource amount is used to characterize the resource quantity information based on which a hyperparameter is verified. A first calculation unit, configured to calculate the full-resource verification results to obtain a first calculation value when the resource amount is all resources. A second calculation unit, configured to calculate the partial-resource verification results to obtain a second calculation value when the resource amount is partial resources. A first determination unit, configured to determine the first calculation value or the second calculation value as the verification result calculation value.
12. An apparatus for processing hyperparameters according to claim 9, wherein, The hyperparameter to be verified generation unit includes: A first update unit configured to update the current integrated probability surrogate model based on the historical verification results and the imputation results of the target hyperparameters to obtain an updated integrated probability surrogate model; A first generation unit configured to generate the hyperparameters to be verified based on the updated integrated probability surrogate model.
13. A hyperparameter processing device according to claim 12, wherein, the device includes: An initial model acquisition unit configured to acquire an initial full-resource surrogate model and at least one initial partial-resource surrogate model; A weight assignment unit configured to assign weights to the initial full-resource surrogate model and the at least one initial partial-resource surrogate model; wherein the sum of the weights corresponding to the initial full-resource surrogate model and the weights corresponding to the at least one initial partial-resource surrogate model is 1; A second generation unit configured to generate the integrated probability surrogate model based on the initial full-resource surrogate model, the weight of the initial full-resource surrogate model, the at least one initial partial-resource surrogate model, and the weights of the at least one initial partial-resource surrogate model.
14. A hyperparameter processing device according to claim 13, wherein, the first update unit includes: A second update unit configured to update the current full-resource surrogate model in the current integrated probability surrogate model with the full-resource imputation result when the imputation result of the target hyperparameters includes a full-resource imputation result corresponding to the full resources; A third update unit configured to update the current partial-resource surrogate model in the current integrated probability surrogate model with the partial-resource imputation result when the imputation result of the target hyperparameters includes a partial-resource imputation result corresponding to the partial resources; A third generation unit configured to obtain the updated integrated probability surrogate model based on the updated current full-resource surrogate model and the updated current partial-resource surrogate model.
15. A hyperparameter processing device according to claim 13, wherein, the device further includes: A selection unit configured to select at least one set of hyperparameter combinations; the hyperparameter combinations include at least two hyperparameters; A second determination unit configured to determine a first verification result corresponding to the at least one set of hyperparameter combinations; the first verification result is a result obtained by verifying the at least one set of hyperparameter combinations based on the full resources; A third determination unit configured to determine a second verification result corresponding to the at least one set of hyperparameter combinations; the second verification result is a result obtained by verifying the at least one set of hyperparameter combinations based on the partial resources; A consistency check unit configured to perform a consistency check on the partial order relationship of the first verification result and the partial order relationship of the second verification result. A weight adjustment unit, configured to perform weight adjustment of the full - scale resource proxy model and the at least one partial resource proxy model based on the consistency check result.
16. A hyperparameter processing device according to claim 14 or 15, wherein, the device further includes: A first resource amount determination unit, configured to obtain a first resource amount based on the weight of the current full - scale resource proxy model and the resource amount corresponding to the current full - scale resource proxy model; A second resource amount determination unit, configured to obtain a second resource amount based on the weight of the current partial resource proxy model and the resource amount corresponding to the current partial resource proxy model; An initial resource amount determination unit, configured to calculate the initial resource amount used for verifying the hyperparameter to be verified according to the first resource amount and the second resource amount.
17. An electronic device, wherein, it includes: a processor; a memory for storing executable instructions of the processor; wherein, the processor is configured to execute the instructions to implement the hyperparameter processing method according to any one of claims 1 to 8.
18. A computer - readable storage medium, wherein, when the instructions in the computer - readable storage medium are executed by a processor of an electronic device, the electronic device can execute the hyperparameter processing method according to any one of claims 1 to 8.
19. A computer program product, including a computer program / instructions, wherein, the computer program / instructions, when executed by a processor, implement the hyperparameter processing method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Deep learning parallel computing architecture method and hyper-parameter automatic configuration optimization thereof
CN111709519A
Hyper-parameter determination method and device, computer equipment and storage medium
CN112529211A