Hyper-parameter optimization method, system and equipment based on deep learning and storage medium

Through the hyperparameter optimization method based on deep learning, the hyperparameter value is dynamically adjusted, which solves the problems of long and high resource cost of hyperparameter tuning in the existing technology, and achieves faster and more efficient model training.

CN120012881AActive Publication Date: 2025-05-16HANGZHOU BYTE ARK TECH CO LTD
View PDF 12 Cites 0 Cited by

Patent Information

Application Number
CN202510480166.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-05-16
Estimated Expiration
2045-04-17

AI Technical Summary

Technical Problem

Existing hyperparameter tuning methods, such as grid search and random search, lead to long training time and excessive computing resources when facing complex models and large data sets.

Method used

The hyperparameter optimization method based on deep learning is adopted. By generating a set of hyperparameter values, the initial parameter value is obtained, and training is performed based on the preset training model, the similarity between the training curve and the preset convergence curve is analyzed, and the hyperparameter value is dynamically adjusted to improve the speed and efficiency of hyperparameter optimization.

Benefits of technology

It significantly improves the speed of hyperparameter optimization, reduces the model training time and resource cost, and promotes efficient development and rapid iteration of deep learning applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120012881A_ABST
    Figure CN120012881A_ABST
Patent Text Reader

Abstract

The invention relates to a hyper-parameter optimization method, system and device based on deep learning and a storage medium, and the method comprises the following steps: repeatedly obtaining an initial parameter value according to an updated hyper-parameter value set, obtaining a training result, and obtaining a corresponding training curve according to the training result; each time one training result is obtained, whether the training curve is similar to a preset convergence curve is judged, and if the training curve is similar to the preset convergence curve, the initial parameter value serves as an optimization parameter corresponding to the to-be-trained model data; if the training curve is not similar to the preset convergence curve, repeatedly obtaining the adjustment value according to the initial parameter value and the first preset value, updating the initial parameter value according to the adjustment value, and updating the training result according to the updated initial parameter value; until the initial parameter value does not belong to the hyper-parameter value set. According to the method, the model training time is shortened, the resource cost is reduced, and efficient development and rapid iteration of deep learning application are effectively promoted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of hyperparameter calculation, and in particular to a hyperparameter optimization method, system, device and storage medium based on deep learning. Background Art

[0002] Hyperparameters are parameters set before the machine learning process, rather than parameter data obtained through training. Usually, hyperparameters need to be optimized to improve the performance and effect of machine learning.

[0003] Hyperparameter tuning is currently performed using either grid search or random search, specifically, by exhaustively searching a manually specified subset of the learning algorithm's hyperparameter space. Grid search algorithms must be guided by some performance metric, typically measured by cross-validation on the training set or by evaluation on a held-out validation set, while random search simply searches through parameter settings a fixed number of times.

[0004] Existing hyperparameter adjustment methods, such as using grid search or random search for hyperparameter tuning, require multiple adjustments to the randomly generated hyperparameters to achieve the optimal selection when faced with complex models and large data sets. This results in a long model training time and consumes a lot of computing resources. Summary of the invention

[0005] In order to reduce the model training time and computing resource usage and improve the model's rapid training and optimization capabilities, the present application provides a hyperparameter optimization method, system, device and storage medium based on deep learning.

[0006] In the first aspect, the present application provides a hyperparameter optimization method based on deep learning, which adopts the following technical solution: A hyperparameter optimization method based on deep learning, comprising the following steps: Generate a set of hyperparameter values ​​based on the model data to be trained, and repeatedly obtain initial parameter values ​​based on the updated set of hyperparameter values; Performing training based on a preset training model and the updated initial parameter values ​​to obtain a training result, and obtaining a corresponding training curve according to the training result; Each time a training result is obtained, it is determined whether the training curve is similar to a preset convergence curve. If the training curve is similar to the preset convergence curve, the initial parameter value is used as the optimization parameter corresponding to the model data to be trained; If the training curve is not similar to the preset convergence curve, repeatedly obtaining an adjustment value according to the initial parameter value and the first preset value, updating the initial parameter value according to the adjustment value, and updating the training result according to the updated initial parameter value; Until the initial parameter value does not belong to the hyperparameter value set, the hyperparameter value set is updated, and the training result is repeatedly updated according to the updated hyperparameter value set.

[0007] By adopting the above technical solution, corresponding hyperparameter value sets are set for different module data to be trained, the range and type of the hyperparameter value set are defined, the training results are obtained through the initial parameter values, and the training curve corresponding to the training results is analyzed to determine whether it is similar to the preset convergence curve. If the training curve is similar to the preset convergence curve, the initial parameter value is determined to be the optimal parameter. The initial parameter value is used as the hyperparameter corresponding to the model data to be trained, which can significantly improve the speed of hyperparameter optimization, reduce the model training time and resource costs, and effectively promote the efficient development and rapid iteration of deep learning applications.

[0008] In some embodiments, the step of obtaining the adjustment value according to the initial parameter value and the first preset value includes the following steps: Obtaining a training type according to the training curve, and determining whether the training type is in a convergence state; If the training type is determined to be a convergence state, obtaining a corresponding convergence degree according to the training result, updating a first preset value based on the convergence degree and a preset convergence degree, and obtaining the adjustment value based on the initial parameter value and the updated first preset value; If it is determined that the training type is not in a convergence state, a replacement parameter value is obtained according to the initial parameter value and the hyperparameter value set, and the replacement parameter value is used as the adjustment value.

[0009] By adopting the above technical solution, when it is determined that the training type is in a convergence state, the corresponding convergence degree is obtained according to the training result, and the first preset value is updated based on the convergence degree and the preset convergence degree, and the adjustment value is obtained based on the initial parameter value and the updated first preset value; when it is determined that the training type is not in a convergence state, a replacement parameter value is obtained according to the initial parameter value and the hyperparameter value set, and the replacement parameter value is used as the adjustment value, and then the adjustment value can be dynamically optimized, thereby improving the efficiency of obtaining hyperparameters, reducing the model training time and the occupancy of computing resources, and improving the model's rapid training and optimization capabilities.

[0010] In some of the embodiments, after obtaining the adjustment value based on the initial parameter value and the updated first preset value, the following steps are also included: Using the training curve corresponding to the training result as a control data curve, obtaining the corresponding training result according to the adjustment value, and using the training curve corresponding to the training result as the current data curve; Determine whether the current data curve has a better convergence effect than the reference data curve; If the current data curve has a better convergence effect than the control data curve, the sum of the initial parameter value and the updated first preset value is used as the adjustment value; If the current data curve does not converge better than the control data curve, the difference between the initial parameter value and the updated first preset value is used as the adjustment value.

[0011] By adopting the above technical solution, the convergence states of the current data curve and the control data curve are compared, so as to reasonably adjust the dynamic search direction of the hyperparameters, improve the efficiency of hyperparameter acquisition, and make the obtained hyperparameters convenient for improving the overall data analysis.

[0012] In some embodiments, updating the set of hyperparameter values ​​comprises the following steps: Obtaining the first value and the last value of the hyperparameter value set, and determining whether the training result corresponding to the first value converges; If the training result corresponding to the first value converges, then determining whether the training result corresponding to the last value converges; If the training result corresponding to the last value converges, a new set of hyperparameter values ​​is generated according to the last value and the first value; If the training result corresponding to the last value does not converge, a new set of hyperparameter values ​​is generated according to the first value; If the training result corresponding to the first value does not converge, determining whether the training result corresponding to the last value converges; If the training result corresponding to the last value converges, a new set of hyperparameter values ​​is generated according to the last value; If the training result corresponding to the last value does not converge, the data type is obtained according to the model data to be trained, and the corresponding control data is screened in the historical database according to the data type; The hyperparameter set corresponding to the control data is used as the updated hyperparameter value set.

[0013] By adopting the above technical solution, the training results of the first value and the last value of the hyperparameter value set are compared. If the training result corresponding to the first value converges, it is determined whether the training result corresponding to the last value converges; if the training result corresponding to the last value converges, a new hyperparameter value set is generated based on the last value and the first value; if the training result corresponding to the last value does not converge, a new hyperparameter value set is generated based on the first value; if the training result corresponding to the first value does not converge, it is determined whether the training result corresponding to the last value converges; if the training result corresponding to the last value converges, a new hyperparameter value set is generated based on the last value; if the training result corresponding to the last value does not converge, the data type is obtained based on the model data to be trained, and the corresponding control data is screened in the historical database based on the data type, and the hyperparameter set corresponding to the control data is used as the updated hyperparameter value set, and the hyperparameter value set is adjusted in real time according to the training results, shortening the search time for the hyperparameter value set, thereby improving the efficiency of hyperparameter acquisition, and making the obtained hyperparameters convenient for improving overall data analysis.

[0014] In some of the embodiments, after the hyperparameter set corresponding to the control data is used as the updated hyperparameter value set, the following steps are further included: The hyperparameter set corresponding to the control data is divided according to preset intervals to obtain the hyperparameter value sets of intervals of different sizes.

[0015] By adopting the above technical solution, hyperparameter sets of different size intervals are obtained, and a parallel computing framework is used to support multi-threaded or multi-node parallel evaluation of different hyperparameter configurations, significantly shortening the optimal hyperparameter acquisition cycle.

[0016] In some embodiments, the training is performed based on a preset training model and the updated initial parameter values, wherein the method for obtaining the preset training model includes the following steps: A corresponding training model is obtained according to the control data, and the training model is used as a preset training model.

[0017] By adopting the above technical solution, the corresponding training model is obtained based on historical data, and the training model is used as a preset training model to predict the value of candidate hyperparameter configurations, thereby achieving more efficient search path planning.

[0018] In some embodiments, generating a new set of hyperparameter values ​​according to the last value and the first value comprises the following steps: If the training result corresponding to the first value converges with the training result corresponding to the last value, a new set of hyperparameter values ​​is generated according to the first value; If the training result corresponding to the last value converges more than the training result corresponding to the first value, a new set of hyperparameter values ​​is generated according to the last value.

[0019] In the second aspect, the present application provides a hyperparameter optimization system based on deep learning, which adopts the following technical solutions: A hyperparameter optimization system based on deep learning, which executes the hyperparameter optimization method based on deep learning according to the first aspect, comprising: A parameter acquisition module, the parameter acquisition module is used to generate a set of hyperparameter values ​​based on the model data to be trained, and repeatedly obtain initial parameter values ​​based on the updated set of hyperparameter values; A curve generation module, the curve generation module is used to perform training based on a preset training model and the updated initial parameter values ​​to obtain a training result, and obtain a corresponding training curve according to the training result; A data processing module, each time a training result is obtained, the data processing module is used to determine whether the training curve is similar to a preset convergence curve, and if the training curve is similar to the preset convergence curve, the data processing module is further used to use the initial parameter value as the optimization parameter corresponding to the model data to be trained; If the training curve is not similar to the preset convergence curve, the data processing module is further used to repeatedly obtain the adjustment value according to the initial parameter value and the first preset value, update the initial parameter value according to the adjustment value, and update the corresponding training result according to the updated initial parameter value; Until the initial parameter value does not belong to the hyperparameter value set, the hyperparameter value set is updated, and the corresponding training result is repeatedly updated according to the updated hyperparameter value set.

[0020] In a third aspect, the present application provides an electronic device, which adopts the following technical solution: An electronic device, comprising a processor and a memory coupled to each other, wherein the memory stores a computer program that can be run on the processor; When the computer program is executed by the processor, the deep learning-based hyperparameter optimization method described in the first aspect is implemented.

[0021] In a fourth aspect, the present application provides a storage medium, which adopts the following technical solution: A storage medium storing at least one instruction, at least one program, a code set or an instruction set, wherein the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement the hyperparameter optimization method based on deep learning as described in the first aspect.

[0022] In summary, the present application includes at least one of the following beneficial technical effects: 1. Set a corresponding set of hyperparameter values ​​for different module data to be trained, define the range and type of the hyperparameter value set, obtain the training results through the initial parameter values, and analyze the training curve corresponding to the training results to determine whether it is similar to the preset convergence curve. If the training curve is similar to the preset convergence curve, the initial parameter value is determined to be the optimal parameter. The initial parameter value is used as the hyperparameter corresponding to the model data to be trained, which can significantly improve the speed of hyperparameter optimization, reduce the model training time and resource costs, and effectively promote the efficient development and rapid iteration of deep learning applications; 2. When it is determined that the training type is in a convergence state, the corresponding convergence degree is obtained according to the training result, and the first preset value is updated based on the convergence degree and the preset convergence degree, and the adjustment value is obtained based on the initial parameter value and the updated first preset value; when it is determined that the training type is not in a convergence state, a replacement parameter value is obtained according to the initial parameter value and the hyperparameter value set, and the replacement parameter value is used as the adjustment value, and then the adjustment value can be dynamically optimized, thereby improving the efficiency of obtaining hyperparameters, reducing the model training time and the occupancy of computing resources, and improving the model's rapid training and optimization capabilities. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 It is a block diagram of a hyperparameter optimization method based on deep learning provided in an embodiment of the present application; Figure 2 is a block diagram of a method for obtaining an adjustment value provided in an embodiment of the present application; Figure 3 It is a block diagram of a method for updating a set of hyperparameter values ​​provided in an embodiment of the present application; Figure 4 is another method block diagram provided by an embodiment of the present application; Figure 5 Schematic diagram of the structure of a hyperparameter optimization system based on deep learning provided in an embodiment of the present application; Figure 6 It is a structural block diagram of the electronic device provided in this embodiment.

[0024] Explanation of the reference numerals: 10, parameter acquisition module; 20, curve generation module; 30, data processing module; 41, processor; 42, memory; 43, computer program. DETAILED DESCRIPTION

[0025] To more clearly understand the purpose, technical solutions and advantages of the present application, the present application is described and illustrated below in conjunction with the accompanying drawings and embodiments. However, it should be understood by those of ordinary skill in the art that the present application can be implemented without these details. In some cases, in order to avoid unnecessary descriptions that make various aspects of the present application obscure, well-known methods, processes, systems, components and / or circuits that have been described at a higher level will not be described in detail. For those of ordinary skill in the art, it is obvious that various changes can be made to the embodiments disclosed in the present application, and without departing from the principles and scope of the present application, the general principles defined in the present application can be applied to other embodiments and application scenarios. Therefore, the present application is not limited to the embodiments shown, but conforms to the broadest scope consistent with the scope claimed for protection of the present application.

[0026] An embodiment of the present application discloses a hyperparameter optimization method based on deep learning, which is applied to a hyperparameter optimization system based on deep learning. The system includes an electronic device, and a processor in the electronic device processes data according to the hyperparameter optimization method based on deep learning, so as to quickly obtain hyperparameters.

[0027] like Figure 1 As shown in FIG. 1 , the hyperparameter optimization method based on deep learning includes the following steps: S100, generating a set of hyperparameter values ​​based on the model data to be trained, and repeatedly obtaining initial parameter values ​​based on the updated set of hyperparameter values.

[0028] The model data to be trained represents the data for which the hyperparameters need to be obtained. The model data to be trained can be various data that need to be trained at present, for example, the training data for target detection of an autonomous vehicle. The hyperparameter value set represents the learning rate set in advance. The hyperparameter value set is used to screen the hyperparameters corresponding to the model data to be trained. The initial parameter value represents the initial value obtained when performing a hyperparameter search. The initial parameter value belongs to the hyperparameter value set.

[0029] It should be noted that the method for obtaining the data of the model to be trained can be based on manual input, or the processor can directly obtain the data of the model to be trained on the Internet. In addition, as to how to obtain the initial parameter value in the hyperparameter value set, a random method can be adopted, or the corresponding initial parameter value can be obtained according to the type of the model data to be trained. In this embodiment, the initial parameter value obtains the middle value of the hyperparameter value set.

[0030] In addition, since the purpose of this application is to filter out hyperparameters, the model data to be trained is not all data, but is used to filter out data corresponding to the hyperparameters. Of course, the model data to be trained can also be all data.

[0031] S200, performing training based on a preset training model and updated initial parameter values ​​to obtain training results, and obtaining corresponding training curves according to the training results.

[0032] Among them, the preset training model represents a training model for obtaining hyperparameters. The training model trains the model data to be trained multiple times until the obtained training results converge with good convergence, which represents that the corresponding initial parameter values ​​can be used as the optimal hyperparameters. In subsequent processing, the training results of the model data to be trained can be quickly obtained. The preset training model can be a real data training model. As long as the hyperparameters are screened, the corresponding hyperparameters can be obtained. The preset training model here can be a low-complexity model generated by pre-training with historical data, that is, a small neural network or a gradient boosting tree.

[0033] The training curve represents a smooth curve obtained based on the training results. The training results obtained here may be divergent, so the training curve is not just a smooth curve, but a curve that includes the most data points in the training results.

[0034] S300, each time a training result is obtained, it is determined whether the training curve is similar to a preset convergence curve. If the training curve is similar to the preset convergence curve, the initial parameter value is used as the optimization parameter corresponding to the model data to be trained.

[0035] Among them, the preset convergence curve is a curve graph under the convergence state, and the preset convergence curve is a curve under the preset convergence state obtained before data training. When a training result is obtained, it is judged whether the training curve is similar to the preset convergence curve. For the training result corresponding to the first initial parameter value, it is not completely similar to the preset convergence curve. After multiple data iterations, the training results are compared multiple times until the training curve is similar to the preset convergence curve, and the initial parameter value is used as the optimization parameter corresponding to the model data to be trained.

[0036] It should be noted here that if the training curve is determined to be similar to the preset convergence curve, the curve trends, convergence starting points and convergence trend values ​​of the two sets of curves can be compared. When the curve trends, convergence starting points and convergence trend values ​​of the two sets of curves are the same, the training curve is determined to be similar to the preset convergence curve. For the two sets of curves whose curve trends, convergence starting points and convergence trend values ​​are not exactly the same, it can be determined whether the training curve has a better convergence effect than the preset convergence curve. If so, the initial parameter value is used as the optimization parameter corresponding to the model data to be trained. If not, the adjustment value is repeatedly obtained based on the initial parameter value and the first preset value, and the initial parameter value is updated based on the adjustment value, and the training result is updated based on the updated initial parameter value.

[0037] S400, if the training curve is not similar to the preset convergence curve, repeatedly obtaining the adjustment value according to the initial parameter value and the first preset value, updating the initial parameter value according to the adjustment value, and updating the training result according to the updated initial parameter value.

[0038] The first preset value represents a value set in advance, which is specifically the amplitude of adjusting the initial parameter value, thereby obtaining a new adjustment value, and using the adjustment value as the new initial parameter value, thereby updating the initial parameter value according to the adjustment value. Updating the training result according to the updated initial parameter value mentioned here specifically refers to repeatedly executing steps S200-S300 to obtain a new training result.

[0039] Combination Figure 2 In one embodiment, when the adjustment range of the initial parameter value is fixed, for some initial parameter values ​​with poor training results, there is a large amount of calculation, which will result in a long model training time and high resource cost. In order to reduce the model training time, the adjustment value is obtained according to the initial parameter value and the first preset value, including the following steps: S410, obtaining a training type according to the training curve, and determining whether the training type is in a convergence state.

[0040] S420, if the training type is determined to be a convergence state, obtain the corresponding convergence degree according to the training result, update the first preset value based on the convergence degree and the preset convergence degree, and obtain the adjustment value based on the initial parameter value and the updated first preset value.

[0041] S430: If it is determined that the training type is not in a convergence state, a replacement parameter value is obtained according to the initial parameter value and the hyperparameter value set, and the replacement parameter value is used as the adjustment value.

[0042] The training types include convergence state and non-convergence state, but are not limited thereto. Other states are also possible, as long as the regularity state can be obtained based on the model data to be trained. Of course, the preset convergence curve can also be a regularity curve obtained by the model data to be trained.

[0043] The degree of convergence is the convergence effect of the training result corresponding to the current initial parameter value. The degree of convergence is compared with the preset degree of convergence. If the degree of convergence is closer to the preset degree of convergence, the first preset value can be set to a smaller value. In this embodiment, the first preset value can be set to 0.0001, and the first preset value can be selected between 0 and 0.01. If the degree of convergence is far from the preset degree of convergence, the first preset value can be set to a larger value. In this embodiment, the first preset value can be set to 0.008, and the first preset value can be selected between 0 and 0.01. The first preset value here can be set according to different situations, specifically according to the comparison between the degree of convergence and the preset degree of convergence.

[0044] When the training type does not converge, it is necessary to obtain replacement parameter values ​​based on the initial parameter values ​​and the hyperparameter value set. The replacement parameter values ​​can change the direction of change based on the initial parameter values, thereby selecting corresponding values ​​in the hyperparameter value set.

[0045] For example, assume that the hyperparameter value set is set to 0~0.01, and the initial parameter value is set to 0.005. After one round of iterations, the initial parameter value searches toward the right side of the hyperparameter value set, but it is found that the training results generated by the values ​​on the right side do not converge. At this time, you can select a value on the left side of the hyperparameter value set as a replacement parameter value.

[0046] S500, until the initial parameter value does not belong to the hyperparameter value set, update the hyperparameter value set, and repeatedly update the training result according to the updated hyperparameter value set.

[0047] Among them, when the initial parameter value does not belong to the hyperparameter value set, it means that the search in the hyperparameter value set has been completed, but no hyperparameter that meets the requirements has been found. At this time, the hyperparameter value set can be updated, and the training results can be repeatedly updated based on the updated hyperparameter value set.

[0048] In one embodiment, updating a set of hyperparameter values ​​includes the following steps: S510, obtaining the first value and the last value of the hyperparameter value set, and determining whether the training result corresponding to the first value converges.

[0049] S520: If the training result corresponding to the first value converges, determine whether the training result corresponding to the last value converges.

[0050] S530: If the training result corresponding to the last value converges, a new set of hyperparameter values ​​is generated based on the last value and the first value.

[0051] S540: If the training result corresponding to the last value does not converge, a new set of hyperparameter values ​​is generated based on the first value.

[0052] S550: If the training result corresponding to the first value does not converge, determine whether the training result corresponding to the last value converges.

[0053] S560: If the training result corresponding to the last value converges, a new set of hyperparameter values ​​is generated based on the last value.

[0054] S570, if the training result corresponding to the last value does not converge, the data type is obtained according to the model data to be trained, and the corresponding control data is screened in the historical database according to the data type.

[0055] S580, taking the hyperparameter set corresponding to the control data as the updated hyperparameter value set.

[0056] The first value represents the beginning value of the hyperparameter value set, and the last value represents the ending value of the hyperparameter value set. The training result corresponding to the first value converges, and the training result corresponding to the last value converges, indicating that the hyperparameter value set belongs to a set in a convergence state. At this time, it is only necessary to infinitely approach the preset convergence curve. Therefore, it is only necessary to obtain the adjustment value according to the first preset value and the initial parameter value until the initial parameter value does not belong to the hyperparameter value set.

[0057] When the training results corresponding to the first value converge and the training results corresponding to the last value do not converge, a new set of hyperparameter values ​​is generated based on the first value. Specifically, the corresponding set of hyperparameter values ​​is generated based on the first value as the center, and the size of the new set of hyperparameter values ​​is consistent with the original set of hyperparameter values.

[0058] When the training result corresponding to the first value does not converge and the training result corresponding to the last value converges, a new set of hyperparameter values ​​is generated based on the last value. Specifically, the corresponding set of hyperparameter values ​​is generated based on the last value as the center, and the size of the new set of hyperparameter values ​​is consistent with the original set of hyperparameter values.

[0059] When the training results corresponding to the first value do not converge and the training results corresponding to the last value do not converge, the data type is obtained according to the model data to be trained, and the corresponding control data is filtered out in the historical database according to the data type, and the hyperparameter set corresponding to the control data is used as the updated hyperparameter value set.

[0060] It should be noted here that the data stored in the historical database can come from similar experiments, tasks or fields. The historical data should be cleaned, normalized and feature extracted to ensure that its format is consistent with the data of the current task. If the historical data has labels, ensure that the labels are consistent with the label space of the new task or can be mapped.

[0061] Reference Figure 3In one embodiment, generating a new set of hyperparameter values ​​according to the last value and the first value includes the following steps: S531: If the training result corresponding to the first value converges with the training result corresponding to the last value, a new set of hyperparameter values ​​is generated based on the first value.

[0062] S532: If the training result corresponding to the last value converges with the training result corresponding to the first value, a new set of hyperparameter values ​​is generated according to the last value.

[0063] Specifically, the new hyperparameter value sets generated are all centered on the first value or the last value and expand on both sides to form a new hyperparameter value set. The size of the formed hyperparameter value set can be the same as the original set or different from the original set. It can be based on the convergence of the training results corresponding to the current hyperparameter value set. If the convergence is good, it can be closer to the original hyperparameter value set. If the convergence is poor, it can be more dispersed with the original hyperparameter value set.

[0064] In one embodiment, after the hyperparameter set corresponding to the control data is used as the updated hyperparameter value set, the following steps are further included: S571, dividing the hyperparameter set corresponding to the control data according to preset intervals to obtain hyperparameter value sets of intervals of different sizes.

[0065] Among them, the hyperparameter value set is divided according to the preset interval, thereby being divided into hyperparameter value sets of different size intervals. A parallel computing framework can be used to evaluate the hyperparameter value set of each different size interval in parallel based on multi-threading or multi-nodes, which can significantly shorten the optimization cycle of the hyperparameters.

[0066] It should be noted here that the specific operation steps are: first, the environment is set up, a multi-node computing cluster is built, and a parallel computing framework (such as MPI, Spark) is configured. Then, the hyperparameter value set is divided into multiple subtasks, that is, each hyperparameter value set of different sizes, each subtask is processed by a computing node or thread, and the subtasks are assigned to each computing node or thread using a task queue or dynamic task allocation strategy. Each computing node or thread evaluates the hyperparameter configuration in parallel and generates an evaluation result. The evaluation results of each computing node or thread are summarized to generate the final hyperparameter optimization result. According to the evaluation results, the hyperparameter configuration is adjusted, and the above steps are repeated until the optimization goal is achieved.

[0067] At the same time, according to the load of the computing nodes, the evaluation tasks of the hyperparameter configuration are dynamically allocated to ensure the load balance of each node. Use task queues (such as Redis and RabbitMQ) to manage the hyperparameter configuration to be evaluated to ensure the orderly allocation and execution of tasks. In distributed computing, design fault-tolerant mechanisms to ensure that when a node fails, tasks can be reallocated to other nodes for continued execution. By dynamically adjusting the task allocation strategy, the load balance of each computing node is ensured to avoid computing bottlenecks. Through asynchronous communication and computing and communication overlap technology, the communication overhead is reduced and the overall computing efficiency is improved. Real-time monitoring of the resource usage of computing nodes (such as CPU, memory, network bandwidth) and timely adjustment of task allocation strategies.

[0068] In one embodiment, training is performed based on a preset training model and updated initial parameter values, wherein the method for obtaining the preset training model includes the following steps: S572, obtaining a corresponding training model based on the control data, and using the training model as a preset training model.

[0069] The control data generated in the historical database is trained using a preset training model, and the set of hyperparameter values ​​corresponding to the control data is used as the new set of hyperparameter values. The preset training model can select a suitable basic model, such as a neural network and a decision tree, which performs well on historical tasks. The model is pre-trained with historical data to learn the features and patterns in historical tasks. The pre-trained model is used as the initial model for new data and fine-tuned on the new data. A smaller learning rate can be used during fine-tuning to avoid destroying useful features learned in the preset training model.

[0070] Specifically, the feature extraction part of the preset training model (such as the first few layers of the convolutional neural network) is directly applied to the new task, and only the classifier or regressor of the new data is trained. Fine-tune the entire preset training model to adapt it to the new data. The number of layers to be fine-tuned can be determined based on the amount of new data. When the amount of data is small, only the last few layers can be fine-tuned. If the historical tasks and the new data are strongly correlated, multi-task learning can be used to train multiple tasks at the same time and share some model parameters.

[0071] Reference Figure 4 In one embodiment, after obtaining the adjustment value based on the initial parameter value and the updated first preset value, the following steps are also included: S421, using the training curve corresponding to the training result as a control data curve, obtaining the corresponding training result according to the adjustment value, and using the training curve corresponding to the training result as the current data curve.

[0072] S422, determining whether the current data curve has a better convergence effect than the control data curve.

[0073] S423: If the current data curve has a better convergence effect than the reference data curve, the sum of the initial parameter value and the updated first preset value is used as the adjustment value.

[0074] S424: If the current data curve does not converge better than the control data curve, the difference between the initial parameter value and the updated first preset value is used as the adjustment value.

[0075] The control data curve represents the training curve corresponding to the training result of the initial parameter value first obtained during the training process, while the current data curve represents the training curve corresponding to the training result of the adjusted value obtained after training iterations.

[0076] Specifically, if the current data curve has a better convergence effect than the reference data curve, the sum of the initial parameter value and the updated first preset value is used as the adjustment value. If the current data curve does not have a better convergence effect than the reference data curve, the difference between the initial parameter value and the updated first preset value is used as the adjustment value. That is, the hyperparameter search continues in the direction corresponding to the hyperparameter value set until the current data curve in the current hyperparameter search direction has a poor convergence effect, at which time the hyperparameter search direction needs to be adjusted.

[0077] An embodiment of the present application also discloses a hyperparameter optimization system based on deep learning, which is used for a hyperparameter optimization method based on deep learning disclosed in the above embodiment.

[0078] like Figure 5 As shown, a hyperparameter optimization system based on deep learning includes a parameter acquisition module 10, a curve generation module 20 connected to the parameter acquisition module 10, and a data processing module 30 connected to the curve generation module 20. The parameter acquisition module 10 is used to generate a hyperparameter value set based on the model data to be trained, and repeatedly obtain the initial parameter value based on the updated hyperparameter value set. The curve generation module 20 is used to perform training based on the preset training model and the updated initial parameter value to obtain the training result, and obtain the corresponding training curve based on the training result. Each time a training result is obtained, the data processing module 30 is used to determine whether the training curve is similar to the preset convergence curve. If the training curve is similar to the preset convergence curve, the data processing module 30 is also used to use the initial parameter value as the optimization parameter corresponding to the model data to be trained. If the training curve is not similar to the preset convergence curve, the data processing module 30 is also used to repeatedly obtain the adjustment value based on the initial parameter value and the first preset value, and update the initial parameter value based on the adjustment value, and update the corresponding training result according to the updated initial parameter value. Until the initial parameter value does not belong to the hyperparameter value set, the hyperparameter value set is updated, and the corresponding training results are repeatedly updated according to the updated hyperparameter value set.

[0079] The other functions performed in the above-mentioned parameter acquisition module 10, curve generation module 20, and data processing module 30 as well as the technical details of each function are the same as or similar to the corresponding features in the deep learning-based hyperparameter optimization method described above, and therefore will not be repeated here.

[0080] The embodiment of the present application also discloses an electronic device.

[0081] Reference Figure 6 The electronic device includes a processor and a memory coupled to each other, and the memory stores a computer program that can be run on the processor.

[0082] When the computer program is executed by a processor, the hyperparameter optimization method based on deep learning disclosed in the above embodiment is implemented.

[0083] The memory 42 may be a ROM or other type of static storage device capable of storing static information and instructions, a random access memory, or other type of dynamic storage device capable of storing information and instructions, or an electrically erasable programmable read-only memory, a read-only optical disc or other optical disc storage, an optical disc storage (including a compact disc, a laser disc, an optical disc, a digital versatile disc, a Blu-ray disc, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium capable of carrying or storing desired program codes in the form of instructions or data structures and capable of being accessed by a computer, but not limited thereto. The memory 42 may be an internal storage unit in some embodiments.

[0084] The processor 41 may be a central processing unit, a general purpose processor, a digital signal processor, an application specific integrated circuit, a field programmable gate array or other programmable logic devices, transistor logic devices, hardware components or any combination thereof, for running program codes stored in the memory 42 or processing data.

[0085] The processor 41 and the memory 42 are connected via a bus. The bus may include a path to transmit information between the above components. The bus may be a peripheral component interconnect standard bus or an extended industrial standard structure bus. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 6 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.

[0086] Figure 6 Only an electronic device having a memory 42, a processor 41, and a bus is shown, and those skilled in the art can understand that Figure 6The structure shown does not constitute a limitation on the electronic device, which can be a bus structure or a star structure. The electronic device can also include more or fewer components than shown in the figure, or combine certain components, or deploy different components. Other existing or future electronic devices may be applicable, should also be included in the scope of protection, and are included here by reference.

[0087] An embodiment of the present application also discloses a storage medium, which stores at least one instruction, at least one program, code set or instruction set, and the at least one instruction, at least one program, code set or instruction set is loaded and executed by a processor to implement the deep learning-based hyperparameter optimization method disclosed in the above embodiment.

[0088] The implementation principle is: First, the parameter acquisition module 10 generates a set of hyperparameter values ​​based on the model data to be trained, and repeatedly obtains initial parameter values ​​based on the updated hyperparameter value set. The curve generation module 20 performs training based on the preset training model and the updated initial parameter values ​​to obtain training results, and obtains the corresponding training curve based on the training results.

[0089] Next, each time a training result is obtained, the data processing module 30 determines whether the training curve is similar to the preset convergence curve. If the training curve is similar to the preset convergence curve, the initial parameter value is used as the optimization parameter corresponding to the model data to be trained.

[0090] Finally, if the training curve is not similar to the preset convergence curve, the data processing module 30 repeatedly obtains the adjustment value based on the initial parameter value and the first preset value, updates the initial parameter value based on the adjustment value, and updates the training result based on the updated initial parameter value until the initial parameter value does not belong to the hyperparameter value set, updates the hyperparameter value set, and repeatedly updates the training result based on the updated hyperparameter value set.

[0091] It should be understood that although the steps in the flowchart of the accompanying drawings are shown in sequence according to the instructions of the arrows, these steps are not necessarily performed in the order indicated by the arrows. Unless otherwise clearly stated in this document, the execution of these steps is not strictly limited in order and can be performed in other orders.

[0092] The above are all preferred embodiments of the present application, and the protection scope of the present application is not limited thereto. Therefore, any equivalent changes made according to the structure, shape, and principle of the present application should be included in the protection scope of the present application.

Claims

1. A hyperparameter optimization method based on deep learning, characterized in that: The following steps are involved: Generate a set of hyperparameter values ​​based on the model data to be trained, and repeatedly obtain initial parameter values ​​based on the updated set of hyperparameter values; Performing training based on a preset training model and the updated initial parameter values ​​to obtain a training result, and obtaining a corresponding training curve according to the training result; Each time a training result is obtained, it is determined whether the training curve is similar to a preset convergence curve. If the training curve is similar to the preset convergence curve, the initial parameter value is used as the optimization parameter corresponding to the model data to be trained; If the training curve is not similar to the preset convergence curve, repeatedly obtaining an adjustment value according to the initial parameter value and the first preset value, updating the initial parameter value according to the adjustment value, and updating the training result according to the updated initial parameter value; Until the initial parameter value does not belong to the hyperparameter value set, the hyperparameter value set is updated, and the training result is repeatedly updated according to the updated hyperparameter value set.

2. The hyperparameter optimization method based on deep learning according to claim 1, characterized in that: The step of obtaining the adjustment value according to the initial parameter value and the first preset value comprises the following steps: Obtaining a training type according to the training curve, and determining whether the training type is in a convergence state; If the training type is determined to be a convergence state, obtaining a corresponding convergence degree according to the training result, updating a first preset value based on the convergence degree and a preset convergence degree, and obtaining the adjustment value based on the initial parameter value and the updated first preset value; If it is determined that the training type is not in a convergence state, a replacement parameter value is obtained according to the initial parameter value and the hyperparameter value set, and the replacement parameter value is used as the adjustment value.

3. The hyperparameter optimization method based on deep learning according to claim 2, characterized in that: After acquiring the adjustment value based on the initial parameter value and the updated first preset value, the following steps are also included: Using the training curve corresponding to the training result as a control data curve, obtaining the corresponding training result according to the adjustment value, and using the training curve corresponding to the training result as the current data curve; Determine whether the current data curve has a better convergence effect than the reference data curve; If the current data curve has a better convergence effect than the control data curve, the sum of the initial parameter value and the updated first preset value is used as the adjustment value; If the current data curve does not converge better than the control data curve, the difference between the initial parameter value and the updated first preset value is used as the adjustment value.

4. The hyperparameter optimization method based on deep learning according to claim 1, characterized in that: The updating of the hyperparameter value set comprises the following steps: Obtaining the first value and the last value of the hyperparameter value set, and determining whether the training result corresponding to the first value converges; If the training result corresponding to the first value converges, then determining whether the training result corresponding to the last value converges; If the training result corresponding to the last value converges, a new set of hyperparameter values ​​is generated according to the last value and the first value; If the training result corresponding to the last value does not converge, a new set of hyperparameter values ​​is generated according to the first value; If the training result corresponding to the first value does not converge, determining whether the training result corresponding to the last value converges; If the training result corresponding to the last value converges, a new set of hyperparameter values ​​is generated according to the last value; If the training result corresponding to the last value does not converge, the data type is obtained according to the model data to be trained, and the corresponding control data is screened in the historical database according to the data type; The hyperparameter set corresponding to the control data is used as the updated hyperparameter value set.

5. The hyperparameter optimization method based on deep learning according to claim 4, characterized in that: After the hyperparameter set corresponding to the control data is used as the updated hyperparameter value set, the following steps are also included: The hyperparameter set corresponding to the control data is divided according to preset intervals to obtain the hyperparameter value sets of intervals of different sizes.

6. The hyperparameter optimization method based on deep learning according to claim 4, characterized in that: The training is performed based on a preset training model and the updated initial parameter values, wherein the method for obtaining the preset training model includes the following steps: A corresponding training model is obtained according to the control data, and the training model is used as a preset training model.

7. The hyperparameter optimization method based on deep learning according to claim 4, characterized in that: The step of generating a new set of hyperparameter values ​​according to the last value and the first value comprises the following steps: If the training result corresponding to the first value converges with the training result corresponding to the last value, a new set of hyperparameter values ​​is generated according to the first value; If the training result corresponding to the last value converges more than the training result corresponding to the first value, a new set of hyperparameter values ​​is generated according to the last value.

8. A hyperparameter optimization system based on deep learning, characterized in that: Executing a hyperparameter optimization method based on deep learning as described in any one of claims 1 to 7, comprising: A parameter acquisition module (10), the parameter acquisition module (10) being used to generate a set of hyperparameter values ​​based on the model data to be trained, and repeatedly acquiring initial parameter values ​​based on the updated set of hyperparameter values; A curve generating module (20), the curve generating module (20) being used to perform training based on a preset training model and the updated initial parameter value to obtain a training result, and to obtain a corresponding training curve based on the training result; A data processing module (30), each time a training result is obtained, the data processing module (30) is used to determine whether the training curve is similar to a preset convergence curve; if the training curve is similar to the preset convergence curve, the data processing module (30) is further used to use the initial parameter value as an optimization parameter corresponding to the model data to be trained; If the training curve is not similar to the preset convergence curve, the data processing module (30) is further used to repeatedly obtain an adjustment value based on the initial parameter value and the first preset value, update the initial parameter value based on the adjustment value, and update the corresponding training result based on the updated initial parameter value; Until the initial parameter value does not belong to the hyperparameter value set, the hyperparameter value set is updated, and the corresponding training result is repeatedly updated according to the updated hyperparameter value set.

9. An electronic device, characterized in that: The electronic device comprises a processor (41) and a memory (42) coupled to each other, wherein the memory (42) stores a computer program (43) that can be run on the processor (41); When the computer program (43) is executed by the processor (41), the deep learning-based hyperparameter optimization method as described in any one of claims 1 to 7 is implemented.

10. A storage medium, characterized in that: The storage medium stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor (41) to implement the hyperparameter optimization method based on deep learning as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Hyper-parameter threshold range determination method and device, storage medium and electronic equipment

    CN109740113A

  • Deep learning chip packaging crack defect detection method based on YOLO

    CN112967243A

  • Deep reinforcement learning model training method and device based on hyper-parameter optimization

    CN113723615A

  • Data processing method, device and equipment and computer readable storage medium

    CN113762514A

  • Operation control method, device and equipment of air conditioner and storage medium

    CN119509011A