Hyperparameter Optimization Method, System, Device and Storage Medium Based on Deep Learning
Through the hyperparameter optimization method based on deep learning, the hyperparameter value set is dynamically adjusted, which solves the problems of long training time and large resource consumption in the existing technology, and achieves fast and efficient hyperparameter optimization.
Patent Information
- Application Number
- CN202510480166.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-04-17
AI Technical Summary
Existing hyperparameter tuning methods such as grid search and random search require multiple adjustments when facing complex models and large data sets, resulting in a long training time and high computing resources consumption.
Using a hyperparameter optimization method based on deep learning, the hyperparameter value is generated, the initial parameter value is obtained and trained, the similarity between the training curve and the preset convergence curve is judged, and the parameter value is dynamically adjusted until the optimal parameter is found.
It significantly improves the speed of hyperparameter optimization, reduces model training time and resource costs, and promotes efficient development and rapid iteration of deep learning.
Smart Images

Figure CN120012881B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of hyperparameter calculation, and in particular, to a hyperparameter optimization method, system, device and storage medium based on deep learning. Background Art
[0002] Hyperparameters are parameters set before the machine learning process, rather than parameter data obtained through training. Generally, hyperparameters need to be optimized to improve the performance and effect of machine learning.
[0003] Currently, grid search or random search is used for hyperparameter tuning. Specifically, it is an exhaustive search of a manually specified subset of the hyperparameter space of the learning algorithm. The grid search algorithm must be guided by certain performance metrics, usually measured by cross-validation on the training set or evaluation on the held-out validation set, while random search is simply a fixed number of searches of the parameter settings.
[0004] Existing hyperparameter adjustment methods, such as using grid search or random search for hyperparameter tuning, when facing complex models and large datasets, the randomly generated hyperparameters need to be adjusted multiple times to reach the optimal choice, resulting in a long model training time and consuming a large amount of computing resources. Summary of the Invention
[0005] In order to reduce the model training time and the consumption of computing resources, and improve the fast training and optimization ability of the model, the present application provides a hyperparameter optimization method, system, device and storage medium based on deep learning.
[0006] In a first aspect, a hyperparameter optimization method based on deep learning provided by the present application adopts the following technical solution:
[0007] A hyperparameter optimization method based on deep learning includes the following steps:
[0008] Generate a set of hyperparameter values based on the data of the model to be trained, and repeatedly obtain the initial parameter values according to the updated set of hyperparameter values;
[0009] Train based on a preset training model and the updated initial parameter values to obtain a training result, and obtain a corresponding training curve according to the training result;
[0010] For each obtained training result, determine whether the training curve is similar to a preset convergence curve. If the training curve is similar to the preset convergence curve, use the initial parameter values as the optimized parameters corresponding to the data of the model to be trained;
[0011] If the training curve is not similar to the preset convergence curve, repeat obtaining an adjustment value based on the initial parameter value and a first preset value, updating the initial parameter value based on the adjustment value, and updating the training result according to the updated initial parameter value;
[0012] Until the initial parameter value does not belong to the hyperparameter value set, update the hyperparameter value set, and repeat updating the training result according to the updated hyperparameter value set.
[0013] By adopting the above technical solution, corresponding hyperparameter value sets are set for different data of the module to be trained, the range and type of the hyperparameter value set are defined, the training result obtained through the initial parameter value is analyzed, and it is judged whether it is similar to the preset convergence curve. If the training curve is similar to the preset convergence curve, it is determined that the initial parameter value is the optimal parameter, and the initial parameter value is used as the hyperparameter corresponding to the data of the model to be trained, which can significantly improve the speed of hyperparameter optimization, reduce the model training time and resource cost, and effectively promote the efficient development and rapid iteration of deep learning applications.
[0014] In some of the embodiments, obtaining the adjustment value based on the initial parameter value and the first preset value includes the following steps:
[0015] Obtain a training type based on the training curve, and judge whether the training type is in a convergence state;
[0016] If it is determined that the training type is in a convergence state, obtain the corresponding convergence degree based on the training result, update the first preset value based on the convergence degree and the preset convergence degree, and obtain the adjustment value based on the initial parameter value and the updated first preset value;
[0017] If it is determined that the training type is not in a convergence state, obtain a replacement parameter value based on the initial parameter value and the hyperparameter value set, and use the replacement parameter value as the adjustment value.
[0018] By adopting the above technical solution, when it is determined that the training type is in a convergence state, obtain the corresponding convergence degree based on the training result, update the first preset value based on the convergence degree and the preset convergence degree, and obtain the adjustment value based on the initial parameter value and the updated first preset value; when it is determined that the training type is not in a convergence state, obtain a replacement parameter value based on the initial parameter value and the hyperparameter value set, and use the replacement parameter value as the adjustment value, so as to dynamically optimize the adjustment value, thereby improving the acquisition efficiency of hyperparameters, reducing the model training duration and the occupation of computing resources, and improving the rapid training and optimization ability of the model.
[0019] In some of these embodiments, after obtaining the adjustment value based on the initial parameter value and the updated first preset value, the following steps are further included:
[0020] Taking the training curve corresponding to the training result as a control data curve, obtaining the corresponding training result according to the adjustment value, and taking the training curve corresponding to the training result as the current data curve;
[0021] Judging whether the convergence effect of the current data curve is better than that of the control data curve;
[0022] If the convergence effect of the current data curve is better than that of the control data curve, then using the sum of the initial parameter value and the updated first preset value as the adjustment value;
[0023] If the convergence effect of the current data curve is not better than that of the control data curve, then using the difference between the initial parameter value and the updated first preset value as the adjustment value.
[0024] By adopting the above technical solution, comparing the convergence states of the current data curve and the control data curve, thereby reasonably adjusting the dynamic search direction of the hyperparameters, improving the efficiency of obtaining the hyperparameters, and enabling the obtained hyperparameters to facilitate improving the overall data analysis.
[0025] In some of these embodiments, updating the hyperparameter value set includes the following steps:
[0026] Obtaining the first value and the last value of the hyperparameter value set, and judging whether the training result corresponding to the first value converges;
[0027] If the training result corresponding to the first value converges, then judging whether the training result corresponding to the last value converges;
[0028] If the training result corresponding to the last value converges, then generating a new hyperparameter value set according to the last value and the first value;
[0029] If the training result corresponding to the last value does not converge, then generating a new hyperparameter value set according to the first value;
[0030] If the training result corresponding to the first value does not converge, then judging whether the training result corresponding to the last value converges;
[0031] If the training result corresponding to the last value converges, then generating a new hyperparameter value set according to the last value;
[0032] If the training result corresponding to the last value does not converge, obtain the data type based on the data of the model to be trained, and filter the corresponding control data in the historical database according to the data type;
[0033] Use the set of hyperparameters corresponding to the control data as the updated set of hyperparameter values.
[0034] By adopting the above technical solution, compare the training results of the first value and the last value in the set of hyperparameter values. If the training result corresponding to the first value converges, determine whether the training result corresponding to the last value converges; if the training result corresponding to the last value converges, generate a new set of hyperparameter values based on the last value and the first value; if the training result corresponding to the last value does not converge, generate a new set of hyperparameter values based on the first value; if the training result corresponding to the first value does not converge, determine whether the training result corresponding to the last value converges; if the training result corresponding to the last value converges, generate a new set of hyperparameter values based on the last value; if the training result corresponding to the last value does not converge, obtain the data type based on the data of the model to be trained, and filter the corresponding control data in the historical database according to the data type, and use the set of hyperparameters corresponding to the control data as the updated set of hyperparameter values. Adjust the set of hyperparameter values in real time according to the training results, shorten the search time for the set of hyperparameter values, and thus can improve the efficiency of hyperparameter acquisition, and can make the obtained hyperparameters facilitate improving the overall data analysis.
[0035] In some of the embodiments, after using the set of hyperparameters corresponding to the control data as the updated set of hyperparameter values, the following steps are further included:
[0036] Divide the set of hyperparameters corresponding to the control data according to a preset interval to obtain sets of hyperparameter values in different size intervals.
[0037] By adopting the above technical solution, obtain sets of hyperparameters in different size intervals, and adopt a parallel computing framework to support multi-threaded or multi-node parallel evaluation of different hyperparameter configurations, significantly shortening the optimal hyperparameter acquisition period.
[0038] In some of the embodiments, training is performed based on the preset training model and the updated initial parameter values, wherein the obtaining method of the preset training model includes the following steps:
[0039] Obtain the corresponding training model according to the control data, and use the training model as the preset training model.
[0040] By adopting the above technical solution, obtain the corresponding training model based on historical data, and use the training model as the preset training model to predict the value of candidate hyperparameter configurations, realizing a more efficient search path planning.
[0041] In some of these embodiments, generating a new set of hyperparameter values based on the end value and the start value includes the following steps:
[0042] If the training result corresponding to the start value converges faster than the training result corresponding to the end value, generate a new set of the hyperparameter values based on the start value;
[0043] If the training result corresponding to the end value converges faster than the training result corresponding to the start value, generate a new set of the hyperparameter values based on the end value.
[0044] In a second aspect, the present application provides a hyperparameter optimization system based on deep learning, adopting the following technical solution:
[0045] A hyperparameter optimization system based on deep learning, which executes the hyperparameter optimization method based on deep learning described in the first aspect, includes:
[0046] A parameter acquisition module, which is used to generate a set of hyperparameter values based on the data of the model to be trained and repeatedly obtain the initial parameter values according to the updated set of the hyperparameter values;
[0047] A curve generation module, which is used to train based on a preset training model and the updated initial parameter values to obtain a training result, and obtain a corresponding training curve according to the training result;
[0048] A data processing module. Every time a training result is obtained, the data processing module is used to determine whether the training curve is similar to a preset convergence curve. If the training curve is similar to the preset convergence curve, the data processing module is further used to use the initial parameter values as the optimized parameter values corresponding to the data of the model to be trained;
[0049] If the training curve is not similar to the preset convergence curve, the data processing module is further used to repeatedly obtain an adjustment value according to the initial parameter values and a first preset value, update the initial parameter values according to the adjustment value, and update the corresponding training result according to the updated initial parameter values;
[0050] Until the initial parameter values do not belong to the set of the hyperparameter values, update the set of the hyperparameter values, and repeatedly update the corresponding training result according to the updated set of the hyperparameter values.
[0051] In a third aspect, the present application provides an electronic device, adopting the following technical solution:
[0052] An electronic device, the electronic device includes a processor and a memory coupled to each other, and a computer program capable of running on the processor is stored on the memory;
[0053] When the computer program is executed by the processor, it implements the hyperparameter optimization method based on deep learning described in the first aspect.
[0054] In a fourth aspect, the present application provides a storage medium, adopting the following technical solution:
[0055] A storage medium stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set are loaded and executed by a processor to implement the hyperparameter optimization method based on deep learning described in the first aspect.
[0056] In summary, the present application includes at least one of the following beneficial technical effects:
[0057] 1. Set corresponding hyperparameter value sets for different data of modules to be trained, define the range and type of the hyperparameter value sets, obtain the training results through the initial parameter values, and analyze the training curves corresponding to the training results to determine whether they are similar to the preset convergence curves. If the training curve is similar to the preset convergence curve, then determine that the initial parameter value is the optimal parameter, and use the initial parameter value as the hyperparameter corresponding to the data of the model to be trained, which can significantly improve the speed of hyperparameter optimization, reduce both the model training time and resource costs, and effectively promote the efficient development and rapid iteration of deep learning applications;
[0058] 2. When it is determined that the training type is in a convergence state, obtain the corresponding convergence degree based on the training result, update the first preset value based on the convergence degree and the preset convergence degree, and obtain the adjustment value based on the initial parameter value and the updated first preset value; when it is determined that the training type is not in a convergence state, obtain a replacement parameter value based on the initial parameter value and the hyperparameter value set, and use the replacement parameter value as the adjustment value, thereby being able to dynamically optimize the adjustment value, which can improve the acquisition efficiency of hyperparameters, reduce the model training duration and the occupation of computing resources, and improve the rapid training and optimization ability of the model. Description of the Drawings
[0059] Figure 1 is a block diagram of the hyperparameter optimization method based on deep learning provided by an embodiment of the present application;
[0060] Figure 2 is a block diagram of the method for obtaining the adjustment value provided by an embodiment of the present application;
[0061] Figure 3It is a block diagram of the method for updating the set of hyperparameter values provided by the embodiments of the present application;
[0062] Figure 4 It is another block diagram provided by the embodiments of the present application;
[0063] Figure 5 It is a schematic structural diagram of a hyperparameter optimization system based on deep learning provided by the embodiments of the present application;
[0064] Figure 6 It is a block diagram of the structure of the electronic device provided by this embodiment.
[0065] Explanation of reference numerals: 10, parameter acquisition module; 20, curve generation module; 30, data processing module; 41, processor; 42, memory; 43, computer program. Detailed implementation manners
[0066] To understand the purpose, technical solution and advantages of the present application more clearly, the present application will be described and illustrated below with reference to the drawings and embodiments. However, those of ordinary skill in the art should understand that the present application can be implemented without these details. In some cases, well-known methods, processes, systems, components and / or circuits that have been described at a higher level will not be described in detail to avoid unnecessary description from obscuring various aspects of the present application. For those of ordinary skill in the art, it is obvious that various changes can be made to the disclosed embodiments of the present application, and the general principles defined in the present application can be applied to other embodiments and application scenarios without departing from the principles and scope of the present application. Therefore, the present application is not limited to the shown embodiments, but conforms to the broadest scope consistent with the scope claimed in the present application.
[0067] The embodiments of the present application disclose a hyperparameter optimization method based on deep learning, which is applied to a hyperparameter optimization system based on deep learning. The system includes an electronic device, and the processor in the electronic device processes data according to the hyperparameter optimization method based on deep learning, so as to quickly obtain hyperparameters.
[0068] As Figure 1 shown, the hyperparameter optimization method based on deep learning includes the following steps:
[0069] S100, generating a set of hyperparameter values based on the data of the model to be trained, and repeatedly obtaining the initial parameter values according to the updated set of hyperparameter values.
[0070] Among them, the data representation of the model to be trained needs to obtain hyperparameter data. The data of the model to be trained can be various data that need to be trained currently. For example, the training data for object detection of an autonomous vehicle. The set of hyperparameter values represents the learning rate set in advance, and the set of hyperparameter values is used to screen the hyperparameters corresponding to the data of the model to be trained. The initial parameter value represents the initial value obtained during hyperparameter search, and the initial parameter value belongs to the set of hyperparameter values.
[0071] It should be noted here that the acquisition method of the data of the model to be trained can be based on manual input, or the processor can directly obtain the data of the model to be trained on the Internet. In addition, for how to obtain the initial parameter value in the set of hyperparameter values, it can be obtained randomly, or according to the type of the data of the model to be trained, the corresponding initial parameter value can be obtained. In this embodiment, the initial parameter value obtains the intermediate value of the set of hyperparameter values.
[0072] In addition, since this application is to screen out hyperparameters, the data of the model to be trained is not all data, but the data corresponding to the hyperparameters to be screened. Of course, the data of the model to be trained can also be all data.
[0073] S200, perform training based on the preset training model and the updated initial parameter value to obtain a training result, and obtain the corresponding training curve according to the training result.
[0074] Among them, the preset training model represents the training model for obtaining hyperparameters. This training model performs multiple trainings on the data of the model to be trained until the obtained training result converges and has good convergence, indicating that the corresponding initial parameter value can be used as the optimal hyperparameter. When processing subsequently, the training result of the data of the model to be trained can be obtained quickly. The preset training model can be a data training model that actually exists, as long as the hyperparameters are screened to obtain the corresponding hyperparameters. The preset training model here can be a low-complexity model generated by pre-training with historical data, that is, a small neural network or a gradient boosting tree.
[0075] The training curve represents a smooth curve obtained according to the training result. The training result obtained here may be divergent. Therefore, the training curve is not just a smooth curve, but includes the curve with the most data points in the training result.
[0076] S300, each time a training result is obtained, and it is judged whether the training curve is similar to the preset convergence curve. If the training curve is similar to the preset convergence curve, the initial parameter value is used as the optimized parameter corresponding to the data of the model to be trained.
[0077] Among them, the preset convergence curve is a curve graph in a convergent state, and this preset convergence curve is the curve in the preset convergent state obtained before data training. When a training result is obtained, it is determined whether the training curve is similar to the preset convergence curve. For the training result corresponding to the first initial parameter value, it cannot be completely similar to the preset convergence curve. After multiple data iterations and multiple comparisons of the training results, until the training curve is similar to the preset convergence curve, the initial parameter value is used as the optimized parameter corresponding to the data of the model to be trained.
[0078] It should be noted here that if it is determined that the training curve is similar to the preset convergence curve, the curve trend, convergence starting point, and convergence tendency value of the two groups of curves can be compared here. When the curve trend, convergence starting point, and convergence tendency value of the two groups of curves are the same, it is determined that the training curve is similar to the preset convergence curve. For the curve trend, convergence starting point, and convergence tendency value of the two groups of curves that are not exactly the same, specifically, it can be determined whether the training curve has a better convergence effect than the preset convergence curve. If so, the initial parameter value is used as the optimized parameter corresponding to the data of the model to be trained. If not, the adjustment value is obtained repeatedly according to the initial parameter value and the first preset value, the initial parameter value is updated according to the adjustment value, and the training result is updated according to the updated initial parameter value.
[0079] S400, if the training curve is not similar to the preset convergence curve, the adjustment value is obtained repeatedly according to the initial parameter value and the first preset value, the initial parameter value is updated according to the adjustment value, and the training result is updated according to the updated initial parameter value.
[0080] Among them, the first preset value represents a value set in advance. Specifically, this value is the amplitude for adjusting the initial parameter value, so as to obtain a new adjustment value and use the adjustment value as the new initial parameter value, thereby realizing the update of the initial parameter value according to the adjustment value. Here, updating the training result according to the updated initial parameter value specifically means repeating steps S200 - S300 to obtain a new training result.
[0081] Combined with Figure 2 , in one of the embodiments, when the adjustment amplitude of the initial parameter value is fixed, for some initial parameter values with poor training results, there is a large amount of calculation, so the model training time is long and the resource cost is high. In order to reduce the model training duration, obtaining the adjustment value according to the initial parameter value and the first preset value includes the following steps:
[0082] S410, obtain the training type according to the training curve and determine whether the training type is a convergent state.
[0083] S420, if it is determined that the training type is in a convergent state, obtain the corresponding convergence degree based on the training result, update the first preset value based on the convergence degree and the preset convergence degree, and obtain the adjustment value based on the initial parameter value and the updated first preset value.
[0084] S430, if it is determined that the training type is not in a convergent state, obtain the replacement parameter value based on the initial parameter value and the set of hyperparameter values, and use the replacement parameter value as the adjustment value.
[0085] Among them, the training type includes a convergent state and a non-convergent state. Of course, it is not limited to this, and there can be other states as long as it is a regular state that can be obtained based on the data of the model to be trained. Of course, the above preset convergence curve can also be a regular curve obtained from the data of the model to be trained.
[0086] The convergence degree represents the convergence effect of the training result corresponding to the current initial parameter value. Compare the convergence degree with the preset convergence degree. If the convergence degree is relatively close to the preset convergence degree, the first preset value can be set to a smaller value. In this embodiment, the first preset value can be set to 0.0001, and the first preset value can be selected between 0 and 0.01. If the convergence degree is far from the preset convergence degree, the first preset value can be set to a larger value. In this embodiment, the first preset value can be set to 0.008, and the first preset value can be selected between 0 and 0.01. The first preset value here can be set according to different situations, specifically based on the comparison between the convergence degree and the preset convergence degree.
[0087] When the training type is non-convergent, it is necessary to obtain the replacement parameter value based on the initial parameter value and the set of hyperparameter values. The replacement parameter value can change the direction of change according to the initial parameter value, so as to select the corresponding value in the set of hyperparameter values.
[0088] Exemplarily, assume that the set of hyperparameter values is set to 0 to 0.01, and the initial parameter value is set to 0.005. After one round of iteration, the initial parameter value starts to search towards the right side of the set of hyperparameter values, but it is found that the training result generated by the values on the right side is non-convergent. At this time, a value on the left side of the set of hyperparameter values can be selected as the replacement parameter value.
[0089] S500, until the initial parameter value does not belong to the set of hyperparameter values, update the set of hyperparameter values, and repeat updating the training result based on the updated set of hyperparameter values.
[0090] Among them, when the initial parameter value does not belong to the set of hyperparameter values, it means that the search in the set of hyperparameter values has been completed, but no hyperparameter that meets the requirements has been found. At this time, the set of hyperparameter values can be updated, and the training result can be updated repeatedly based on the updated set of hyperparameter values.
[0091] In one of the embodiments, updating the set of hyperparameter values includes the following steps:
[0092] S510. Obtain the first value and the last value of the set of hyperparameter values, and determine whether the training result corresponding to the first value converges.
[0093] S520. If the training result corresponding to the first value converges, then determine whether the training result corresponding to the last value converges.
[0094] S530. If the training result corresponding to the last value converges, then generate a new set of hyperparameter values based on the last value and the first value.
[0095] S540. If the training result corresponding to the last value does not converge, then generate a new set of hyperparameter values based on the first value.
[0096] S550. If the training result corresponding to the first value does not converge, then determine whether the training result corresponding to the last value converges.
[0097] S560. If the training result corresponding to the last value converges, then generate a new set of hyperparameter values based on the last value.
[0098] S570. If the training result corresponding to the last value does not converge, then obtain the data type based on the data of the model to be trained, and filter the corresponding reference data in the historical database according to the data type.
[0099] S580. Use the set of hyperparameters corresponding to the reference data as the updated set of hyperparameter values.
[0100] Wherein, the first value represents the starting value of the set of hyperparameter values, and the last value represents the ending value of the set of hyperparameter values. The convergence of the training result corresponding to the first value and the convergence of the training result corresponding to the last value indicate that the set of hyperparameter values belongs to a convergent state set. At this time, it only needs to approach the preset convergence curve infinitely. Therefore, only obtain the adjustment value according to the first preset value and the initial parameter value until the initial parameter value does not belong to the set of hyperparameter values.
[0101] When the training result corresponding to the first value converges and the training result corresponding to the last value does not converge, then generate a new set of hyperparameter values based on the first value. Specifically, a corresponding set of hyperparameter values is generated centered on the first value, and the size of the new set of hyperparameter values formed is the same as that of the original set of hyperparameter values.
[0102] When the training result corresponding to the first value does not converge and the training result corresponding to the last value converges, a new set of hyperparameter values is generated based on the last value. Specifically, a corresponding set of hyperparameter values is generated centered around the last value, and the size of the newly formed set of hyperparameter values is the same as that of the original set of hyperparameter values.
[0103] When the training result corresponding to the first value does not converge and the training result corresponding to the last value does not converge, the data type is obtained based on the data to be trained, and the corresponding reference data is filtered from the historical database according to the data type. The set of hyperparameters corresponding to the reference data is used as the updated set of hyperparameter values.
[0104] It should be noted here that the data stored in the historical database can come from similar experiments, tasks, or fields. The historical data is cleaned, normalized, and feature extracted to ensure that its format is consistent with the data of the current task. If the historical data has labels, ensure that the labels are consistent with the label space of the new task or can be mapped.
[0105] Refer to Figure 3 In one of the embodiments, generating a new set of hyperparameter values based on the last value and the first value includes the following steps:
[0106] S531, if the training result corresponding to the first value converges more than the training result corresponding to the last value, a new set of hyperparameter values is generated based on the first value.
[0107] S532, if the training result corresponding to the last value converges more than the training result corresponding to the first value, a new set of hyperparameter values is generated based on the last value.
[0108] Specifically, the newly generated sets of hyperparameter values are centered around the first value or the last value and expand to both sides to form new sets of hyperparameter values. The size of the formed set of hyperparameter values can be the same as or different from the original set, and specifically, it can be determined according to the convergence situation of the training result corresponding to the current set of hyperparameter values. If the convergence situation is better, it can be closer to the original set of hyperparameter values; if the convergence situation is worse, it can be more dispersed from the original set of hyperparameter values.
[0109] In one of the embodiments, after using the set of hyperparameters corresponding to the reference data as the updated set of hyperparameter values, the following steps are further included:
[0110] S571, the set of hyperparameters corresponding to the reference data is segmented according to a preset interval to obtain sets of hyperparameter values with different sizes of intervals.
[0111] Among them, the hyperparameter value set is segmented according to a preset interval, so as to be segmented into hyperparameter value sets with different sizes. A parallel computing framework can be used to parallelly evaluate each hyperparameter value set with a different size according to multiple threads or multiple nodes, significantly shortening the optimization period of hyperparameters.
[0112] It should be noted here that the specific operation steps are as follows. First, the environment is set up, a multi-node computing cluster is built, and a parallel computing framework (such as MPI, Spark) is configured. Then, the hyperparameter value set is segmented into multiple sub-tasks, that is, hyperparameter value sets with different sizes. Each sub-task is processed by a computing node or thread. A task queue or a dynamic task allocation strategy is used to allocate the sub-tasks to each computing node or thread. Each computing node or thread parallelly evaluates the hyperparameter configuration to generate an evaluation result. The evaluation results of each computing node or thread are aggregated to generate the final hyperparameter optimization result. According to the evaluation result, the hyperparameter configuration is adjusted, and the above steps are repeated until the optimization goal is achieved.
[0113] At the same time, according to the load conditions of the computing nodes, the evaluation tasks of the hyperparameter configuration are dynamically allocated to ensure the load balance of each node. A task queue (such as Redis, RabbitMQ) is used to manage the hyperparameter configurations to be evaluated, ensuring the orderly allocation and execution of tasks. In distributed computing, a fault tolerance mechanism is designed to ensure that when a certain node fails, the task can be re-allocated to other nodes for continued execution. By dynamically adjusting the task allocation strategy, the load balance of each computing node is ensured, and computational bottlenecks are avoided. Through asynchronous communication and techniques of overlapping computing and communication, the communication overhead is reduced, and the overall computing efficiency is improved. The resource usage conditions (such as CPU, memory, network bandwidth) of the computing nodes are monitored in real time, and the task allocation strategy is adjusted in a timely manner.
[0114] In one of the embodiments, training is performed based on a preset training model and the updated initial parameter values. Among them, the acquisition method of the preset training model includes the following steps:
[0115] S572, obtain the corresponding training model according to the control data, and use the training model as the preset training model.
[0116] The control data generated in the historical database is trained using a preset training model, and the hyperparameter value set corresponding to the control data is used as the new hyperparameter value set. A suitable basic model can be selected for the preset training model, such as a neural network and a decision tree, etc. This model performs well on historical tasks. The historical data is used to pre-train the model so that it learns the features and patterns in the historical tasks. The pre-trained model is used as the initial model for the new data and is fine-tuned on the new data. A smaller learning rate can be used during fine-tuning to avoid destroying the useful features learned by the preset training model.
[0117] Specifically, the feature extraction part of the preset training model (such as the first few layers of a convolutional neural network) is directly applied to the new task, and only the classifier or regressor for the new data is trained. The entire preset training model is fine-tuned to adapt to the new data. The number of layers to be fine-tuned can be determined according to the amount of new data. When the amount of data is small, only the last few layers can be fine-tuned. If there is a strong correlation between the historical task and the new data, multi-task learning can be adopted to train multiple tasks simultaneously and share some model parameters.
[0118] Refer to Figure 4 , in one of the embodiments, after obtaining the adjustment value based on the initial parameter value and the updated first preset value, the following steps are further included:
[0119] S421, taking the training curve corresponding to the training result as the control data curve, obtaining the corresponding training result according to the adjustment value, and taking the training curve corresponding to the training result as the current data curve.
[0120] S422, judging whether the convergence effect of the current data curve is better than that of the control data curve.
[0121] S423, if the convergence effect of the current data curve is better than that of the control data curve, then using the sum of the initial parameter value and the updated first preset value as the adjustment value.
[0122] S424, if the convergence effect of the current data curve is not better than that of the control data curve, then using the difference between the initial parameter value and the updated first preset value as the adjustment value.
[0123] Among them, the control data curve represents the training curve corresponding to the training result of the initial parameter value obtained first during the training process, while the current data curve is the training curve corresponding to the training result of the adjustment value obtained after training iteration.
[0124] Specifically, if the convergence effect of the current data curve is better than that of the control data curve, then using the sum of the initial parameter value and the updated first preset value as the adjustment value; if the convergence effect of the current data curve is not better than that of the control data curve, then using the difference between the initial parameter value and the updated first preset value as the adjustment value. That is, continue to search for hyperparameters in the direction corresponding to the hyperparameter value set until the convergence effect of the current data curve appearing in the current hyperparameter search direction is poor, and at this time, the hyperparameter search direction needs to be adjusted.
[0125] The embodiment of the present application also discloses a hyperparameter optimization system based on deep learning, which is used for the hyperparameter optimization method based on deep learning disclosed in the above embodiment.
[0126] Such as Figure 5As shown in the figure, a hyperparameter optimization system based on deep learning includes a parameter acquisition module 10, a curve generation module 20 connected to the parameter acquisition module 10 through a network, and a data processing module 30 connected to the curve generation module 20 through a network. The parameter acquisition module 10 is used to generate a set of hyperparameter values based on the data of the model to be trained, and repeatedly obtain the initial parameter values according to the updated set of hyperparameter values. The curve generation module 20 is used to train based on a preset training model and the updated initial parameter values to obtain a training result, and obtain the corresponding training curve according to the training result. For each obtained training result, the data processing module 30 is used to determine whether the training curve is similar to a preset convergence curve. If the training curve is similar to the preset convergence curve, the data processing module 30 is further used to use the initial parameter values as the optimized parameters corresponding to the data of the model to be trained. If the training curve is not similar to the preset convergence curve, the data processing module 30 is further used to repeatedly obtain an adjustment value according to the initial parameter values and a first preset value, update the initial parameter values according to the adjustment value, and update the corresponding training result according to the updated initial parameter values. Until the initial parameter values do not belong to the set of hyperparameter values, update the set of hyperparameter values, and repeatedly update the corresponding training result according to the updated set of hyperparameter values.
[0127] Other functions performed by the above-mentioned parameter acquisition module 10, curve generation module 20, and data processing module 30, as well as the technical details of each function, are the same as or similar to the corresponding features in the previously described hyperparameter optimization method based on deep learning, so they will not be elaborated here.
[0128] The embodiment of the present application also discloses an electronic device.
[0129] Refer to Figure 6 , the electronic device includes a processor and a memory coupled to each other, and a computer program capable of running on the processor is stored on the memory.
[0130] When the computer program is executed by the processor, it implements the hyperparameter optimization method based on deep learning disclosed in the above embodiment.
[0131] The memory 42 can be a ROM or other types of static storage devices that can store static information and instructions, a random access memory, or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory, a compact disc read-only memory, or other optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 42 can be an internal storage unit in some embodiments.
[0132] The processor 41 can be a central processing unit, a general-purpose processor, a data signal processor, an application-specific integrated circuit, a field-programmable gate array or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It is used to run the program code stored in the memory 42 or process data.
[0133] The processor 41 and the memory 42 are connected by a bus. The bus can include a path for transmitting information between the above components. The bus can be a peripheral component interconnect standard bus or an extended industry standard architecture bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity, Figure 6 only a thick line is shown in the figure, but it does not mean that there is only one bus or one type of bus.
[0134] Figure 6 Only an electronic device having a memory 42, a processor 41, and a bus is shown. Those skilled in the art can understand that Figure 6 the shown structure does not constitute a limitation on the electronic device. It can be a bus-type structure or a star structure. The electronic device can also include more or fewer components than those shown in the figure, or combine some components, or have different component deployments. Other existing or future possible electronic devices are applicable and should also be included in the protection scope and are hereby incorporated by reference.
[0135] The embodiment of the present application also discloses a storage medium. The storage medium stores at least one instruction, at least one segment of program, a code set, or an instruction set. The at least one instruction, at least one segment of program, the code set, or the instruction set is loaded and executed by the processor to implement the hyperparameter optimization method based on deep learning disclosed in the above embodiment.
[0136] The implementation principle is as follows:
[0137] First, the parameter acquisition module 10 generates a set of hyperparameter values based on the data of the model to be trained, and repeatedly obtains the initial parameter values according to the updated set of hyperparameter values. The curve generation module 20 trains based on the preset training model and the updated initial parameter values to obtain a training result, and obtains the corresponding training curve according to the training result.
[0138] Next, for each obtained training result, the data processing module 30 determines whether the training curve is similar to the preset convergence curve. If the training curve is similar to the preset convergence curve, the initial parameter values are used as the optimized parameters corresponding to the data of the model to be trained.
[0139] Finally, if the training curve is not similar to the preset convergence curve, the data processing module 30 repeats obtaining an adjustment value according to the initial parameter value and the first preset value, updates the initial parameter value according to the adjustment value, and updates the training result according to the updated initial parameter value until the initial parameter value does not belong to the hyperparameter value set, updates the hyperparameter value set, and repeats updating the training result according to the updated hyperparameter value set.
[0140] It should be understood that although each step in the flowchart of the accompanying drawings is shown in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, there is no strict order restriction for the execution of these steps, and they can be executed in other orders.
[0141] The above are all preferred embodiments of the present application, and the protection scope of the present application is not limited thereby. Therefore, all equivalent changes made according to the structure, shape, and principle of the present application should be covered within the protection scope of the present application.
Claims
1. A hyperparameter optimization method based on deep learning, characterized in that, Processing by a processor to obtain hyperparameters, including the following steps: Generating a set of hyperparameter values based on the data of the model to be trained, and repeatedly obtaining initial parameter values according to the updated set of hyperparameter values; Training based on a preset training model and the updated initial parameter values to obtain a training result, and obtaining a corresponding training curve according to the training result; Each time a training result is obtained, and it is judged whether the training curve is similar to a preset convergence curve. If the training curve is similar to the preset convergence curve, the initial parameter value is used as the optimized parameter corresponding to the data of the model to be trained; If the training curve is not similar to the preset convergence curve, repeatedly obtaining an adjustment value according to the initial parameter value and a first preset value, updating the initial parameter value according to the adjustment value, and updating the training result according to the updated initial parameter value; Until the initial parameter value does not belong to the set of hyperparameter values, updating the set of hyperparameter values, and repeatedly updating the training result according to the updated set of hyperparameter values; Among them, the updating of the set of hyperparameter values includes the following steps: Obtaining the first value and the last value of the set of hyperparameter values, and judging whether the training result corresponding to the first value converges; If the training result corresponding to the first value converges, judging whether the training result corresponding to the last value converges; If the training result corresponding to the last value converges, generating a new set of hyperparameter values according to the last value and the first value; If the training result corresponding to the last value does not converge, generating a new set of hyperparameter values according to the first value; If the training result corresponding to the first value does not converge, judging whether the training result corresponding to the last value converges; If the training result corresponding to the last value converges, generating a new set of hyperparameter values according to the last value; If the training result corresponding to the last value does not converge, obtaining the data type according to the data of the model to be trained, and screening the corresponding control data in the historical database according to the data type; Taking the set of hyperparameters corresponding to the control data as the updated set of hyperparameter values; After taking the set of hyperparameters corresponding to the control data as the updated set of hyperparameter values, the following steps are further included: Dividing the set of hyperparameters corresponding to the control data according to a preset interval to obtain sets of hyperparameter values in different size intervals. Among them, an environment is built, a multi-node computing cluster is built, a parallel computing framework is configured, and the set of hyperparameter values is divided according to the preset interval, so as to be divided into sets of hyperparameters in different sizes. The parallel computing framework is used to evaluate each set of hyperparameter values in different sizes in parallel according to multiple threads or multiple nodes.
2. The hyperparameter optimization method based on deep learning according to claim 1, characterized in that The obtaining of the adjustment value according to the initial parameter value and the first preset value includes the following steps: Obtaining the training type according to the training curve, and judging whether the training type is in a convergence state; If it is determined that the training type is in a converged state, obtain the corresponding convergence degree according to the training result, update the first preset value based on the convergence degree and a preset convergence degree, and obtain the adjustment value based on the initial parameter value and the updated first preset value; If it is determined that the training type is not in a converged state, obtain a replacement parameter value according to the initial parameter value and the set of hyperparameter values, and use the replacement parameter value as the adjustment value.
3. The hyperparameter optimization method based on deep learning according to claim 2, wherein After obtaining the adjustment value based on the initial parameter value and the updated first preset value, the following steps are further included: Use the training curve corresponding to the training result as a control data curve, obtain the corresponding training result according to the adjustment value, and use the training curve corresponding to the training result as the current data curve; Judge whether the convergence effect of the current data curve is better than that of the control data curve; If the convergence effect of the current data curve is better than that of the control data curve, use the sum of the initial parameter value and the updated first preset value as the adjustment value; If the convergence effect of the current data curve is not better than that of the control data curve, use the difference between the initial parameter value and the updated first preset value as the adjustment value.
4. The hyperparameter optimization method based on deep learning according to claim 1, wherein Perform training based on a preset training model and the updated initial parameter value, wherein the obtaining method of the preset training model includes the following steps: Obtain a corresponding training model according to the control data, and use the training model as the preset training model.
5. The hyperparameter optimization method based on deep learning according to claim 1, characterized in that The generating a new set of hyperparameter values according to the last value and the first value includes the following steps: If the training result corresponding to the first value is more convergent than the training result corresponding to the last value, generate a new set of the hyperparameter values according to the first value; If the training result corresponding to the last value is more convergent than the training result corresponding to the first value, generate a new set of the hyperparameter values according to the last value.
6. A hyperparameter optimization system based on deep learning, characterized in that, Implementing the method for optimizing hyperparameters based on deep learning according to any one of claims 1-5, includes: A parameter acquisition module (10), the parameter acquisition module (10) is configured to generate a set of hyperparameter values based on data of a model to be trained, and repeatedly obtain an initial parameter value according to the updated set of hyperparameter values; A curve generation module (20), the curve generation module (20) is configured to perform training based on a preset training model and the updated initial parameter value to obtain a training result, and obtain a corresponding training curve according to the training result; A data processing module (30), each time a training result is obtained, the data processing module (30) is configured to judge whether the training curve is similar to a preset convergence curve, if the training curve is similar to the preset convergence curve, the data processing module (30) is further configured to use the initial parameter value as the optimized parameter corresponding to the data of the model to be trained; If the training curve is not similar to the preset convergence curve, the data processing module (30) is further configured to repeatedly obtain an adjustment value according to the initial parameter value and a first preset value, update the initial parameter value according to the adjustment value, and update the corresponding training result according to the updated initial parameter value; Until the initial parameter value does not belong to the hyperparameter value set, update the hyperparameter value set, and repeatedly update the corresponding training result according to the updated hyperparameter value set.
7. An electronic device, characterized in that, The electronic device includes a processor (41) and a memory (42) that are coupled to each other, and a computer program (43) capable of running on the processor (41) is stored on the memory (42); When the computer program (43) is executed by the processor (41), it implements the deep learning-based hyperparameter optimization method according to any one of claims 1-5.
8. A storage medium, characterized in that, The storage medium stores at least one instruction, at least one segment of program, a code set or an instruction set, and the at least one instruction, the at least one segment of program, the code set or the instruction set is loaded and executed by the processor (41) to implement the deep learning-based hyperparameter optimization method according to any one of claims 1-5.
Citation Information
Patent Citations
Deep reinforcement learning model training method and device based on hyper-parameter optimization
CN113723615A
Target prediction method for small sample data
CN119829906A