Training Method and Device for Deep Reinforcement Learning Model Based on Hyperparameter Optimization
Through the hyperparameter optimization method, the hyperparameters of the deep reinforcement learning model are solved, and the existing offline deep reinforcement learning method has a large deviation in the training effect of different data sets is achieved, and a deep reinforcement learning model with higher performance and wider application is achieved.
Patent Information
- Application Number
- CN202011621981.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-31
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2040-12-31
AI Technical Summary
The existing offline deep reinforcement learning methods have large deviations in training effects at different data sets, resulting in poor adaptability and low performance.
By obtaining multiple initial hyperparameter combinations and deep reinforcement learning models, these hyperparameters are used to train the model and filter out the models with good performance. Then, the initial hyperparameter combination is optimized to form the target hyperparameter combination, and finally, these hyperparameters are used to train the target deep reinforcement learning model.
It realizes the deep reinforcement learning model with higher performance, and makes the model adapt to a wider range of application scenarios, improving the stability and adaptability of the training effect.
Smart Images

Figure CN113723615B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technologies, and in particular, to a training method, apparatus, electronic device, storage medium, and computer program product for a deep reinforcement learning model based on hyperparameter optimization. Background Art
[0002] Deep Reinforcement Learning (abbreviated as Deep RL) is a technology that has emerged in recent years. This technology combines two technologies: deep learning and reinforcement learning. Deep RL has the ability to recognize patterns in high-dimensional states in complex systems and output actions based on this. Based on deep reinforcement learning, learning can be carried out by interacting with the environment and continuously trial-and-error summarization. Deep RL is applicable to control, decision-making, and complex system optimization tasks. In fields such as games, autonomous driving control and decision-making, robot control, finance, and industrial system control optimization, Deep RL has huge potential application space. However, since the training of Deep RL requires large-scale interaction with the environment, this condition is not available in most real-world scenarios, which severely restricts the implementation of deep reinforcement learning methods.
[0003] To solve this problem, the related art has proposed Off-line Deep RL technology. However, currently, the training effect of Off-line Deep RL methods varies greatly with different data sets, resulting in problems such as poor adaptability and low performance in the achievable training effect. Summary of the Invention
[0004] The present application provides a training method and apparatus for a deep reinforcement learning model based on hyperparameter optimization.
[0005] According to a first aspect of the present application, there is provided a training method for a deep reinforcement learning model based on hyperparameter optimization, including:
[0006] Obtaining a plurality of initial hyperparameter combinations and a plurality of first deep reinforcement learning models;
[0007] Training the plurality of first deep reinforcement learning models with a plurality of hyperparameters in the initial hyperparameter combinations to obtain training evaluation indicators respectively corresponding to the plurality of first deep reinforcement learning models;
[0008] Selecting a second deep reinforcement learning model from the plurality of first deep reinforcement learning models according to the training evaluation indicators;
[0009] Optimizing the initial hyperparameter combination by using a plurality of target hyperparameters corresponding to the second deep reinforcement learning model to form a target hyperparameter combination; and
[0010] Training the second deep reinforcement learning model by using a plurality of hyperparameters among the target hyperparameter combination to obtain a target deep reinforcement learning model.
[0011] According to a second aspect of the present application, there is provided a training device for a deep reinforcement learning model based on hyperparameter optimization, including:
[0012] A first acquisition module, configured to acquire a plurality of initial hyperparameter combinations and a plurality of first deep reinforcement learning models;
[0013] A first training module, configured to train the plurality of first deep reinforcement learning models by using a plurality of hyperparameters in the initial hyperparameter combination to obtain training evaluation indexes respectively corresponding to the plurality of first deep reinforcement learning models;
[0014] A first screening module, configured to screen out a second deep reinforcement learning model from the plurality of first deep reinforcement learning models according to the training evaluation indexes;
[0015] A first processing module, configured to optimize the initial hyperparameter combination by using a plurality of target hyperparameters corresponding to the second deep reinforcement learning model to form a target hyperparameter combination; and
[0016] A second training module, configured to train the second deep reinforcement learning model by using a plurality of hyperparameters among the target hyperparameter combination to obtain a target deep reinforcement learning model.
[0017] According to a third aspect of the present application, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the training method for a deep reinforcement learning model based on hyperparameter optimization according to the above-mentioned embodiment of one aspect.
[0018] According to a fourth aspect of the present application, there is provided a non-transitory computer-readable storage medium storing computer instructions, on which a computer program is stored, and the computer instructions are used to cause the computer to execute the training method for a deep reinforcement learning model based on hyperparameter optimization according to the above-mentioned embodiment of one aspect.
[0019] According to a fifth aspect of the present application, there is provided a computer program product, and when the computer program is executed by a processor, the training method for a deep reinforcement learning model based on hyperparameter optimization according to the above-mentioned embodiment of one aspect is implemented.
[0020] In the technical solution of the embodiment of the present application, according to a plurality of initial hyperparameter combinations and a plurality of deep reinforcement learning models, the initial hyperparameter combinations are optimized to form target hyperparameter combinations, and then a plurality of hyperparameters in the target hyperparameter combinations are used to train a target deep reinforcement learning model. Thus, by combining hyperparameter optimization and model training to implement the training of the deep reinforcement learning model, not only can a deep reinforcement learning model with higher performance be trained, but also the trained model can be adapted to a wider range of application scenarios.
[0021] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become easily understood through the following description. Description of the Drawings
[0022] The drawings are used to better understand the present solution and do not constitute a limitation to the present application. Among them:
[0023] Figure 1 is a schematic flowchart of a method for training a deep reinforcement learning model based on hyperparameter optimization provided by an embodiment of the present application;
[0024] Figure 2 is a schematic diagram of the principle of model training provided by an embodiment of the present application;
[0025] Figure 3 is a schematic flowchart of a process for optimizing an initial hyperparameter combination provided by an embodiment of the present application;
[0026] Figure 4 is a schematic flowchart of a process for training a second deep reinforcement learning model provided by an embodiment of the present application;
[0027] Figure 5 is a schematic flowchart of a process for screening a second deep reinforcement learning model provided by an embodiment of the present application;
[0028] Figure 6 is a schematic diagram of the principle of another method for training a deep reinforcement learning model based on hyperparameter optimization provided by an embodiment of the present application;
[0029] Figure 7 is a schematic diagram of applying the training method to a thermal power generation system in the industrial field provided by an embodiment of the present application;
[0030] Figure 8 is a schematic structural diagram of a training device for a deep reinforcement learning model based on hyperparameter optimization provided by an embodiment of the present application;
[0031] Figure 9A block diagram of an electronic device for implementing the training method of the deep reinforcement learning model based on hyperparameter optimization according to the embodiments of the present application. Detailed implementation manners
[0032] The following describes exemplary embodiments of the present application with reference to the accompanying drawings. Various details of the embodiments of the present application are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present application. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0033] The following describes the training method and device of the deep reinforcement learning model based on hyperparameter optimization according to the embodiments of the present application with reference to the accompanying drawings.
[0034] It should be noted that the offline deep reinforcement learning algorithm framework usually includes multiple deep neural networks. The optimization of network parameters in the training process requires alternating training and collaborative optimization of multiple networks. In addition to the relevant parameters of the deep neural network (such as the number of network layers, the size of each network layer, the learning rate, the activation function parameters, etc.), the offline deep reinforcement learning algorithm also has some characteristic parameters (such as the discount factor γ, the buffer size of the cache pool, the soft update parameter τ, etc.). Therefore, compared with general machine learning algorithms, the number of hyperparameters of the offline deep reinforcement learning algorithm is larger, and the value range of the hyperparameters is also larger. Secondly, since most reinforcement learning algorithms use an algorithm framework of multi-network collaborative optimization, a large number of hyperparameters are introduced during the alternating process, making there many random factors in the training process, the model has strong randomness and low reproducibility. At the same time, due to the large amount and wide range of hyperparameter data, the model is very sensitive to the high-dimensional and complex hyperparameter space and is easily affected by hyperparameters. The workload of manual hyperparameter tuning for offline deep reinforcement learning is large and it is difficult to cover all parameter spaces. There is a greater need for an algorithm to "carefully select" hyperparameters. In addition, due to the particularity of the Offline Deep RL algorithm, in addition to convergence and reproducibility, the performance of the model will also be affected by hyperparameters. Moreover, since optimization needs to be performed from data, the optimization amplitude of the existing strategy itself is limited, and the role of hyperparameters will be more important.
[0035] Hyperparameters need to be defined before training a machine learning model and cannot be adjusted during the training process. The selection of hyperparameters plays a very important role in the final performance of the model. For example, a large and complex deep neural network may be suitable for processing data from different sources, but too many layers will ultimately lead to the vanishing gradient and inability to train. Another example is that if the learning rate is too large, it will affect the convergence effect of the model, and if it is too small, it will lead to a too slow convergence speed. For the selection of hyperparameters, generally there will be different optimal combinations for different models. Usually, engineers will try to find this set of optimal hyperparameters based on experience or using random methods. There is no set of perfect hyperparameters suitable for all models. Behind the excellent performance of the model is the process of "meticulous selection" of hyperparameters. Tuning parameters is an unavoidable step in model training. Therefore, hyperparameter optimization is very valuable in both scientific research and engineering.
[0036] Figure 1 FIG. is a schematic flow chart of a training method for a deep reinforcement learning model based on hyperparameter optimization provided by an embodiment of the present application.
[0037] It should be noted that the execution subject of the training method for the deep reinforcement learning model based on hyperparameter optimization in the embodiment of the present application can be an electronic device. Specifically, the electronic device can be, but is not limited to, a server or a terminal, and the terminal can be, but is not limited to, a personal computer, a smart phone, an IPAD, etc.
[0038] In the embodiment of the present application, the training method for the deep reinforcement learning model based on hyperparameter optimization is configured in a training device for the deep reinforcement learning model based on hyperparameter optimization as an example. The device can be applied to an electronic device so that the electronic device can execute the training method for the deep reinforcement learning model based on hyperparameter optimization.
[0039] As Figure 1 shown, the training method for the deep reinforcement learning model based on hyperparameter optimization includes the following steps:
[0040] S101, obtain a plurality of initial hyperparameter combinations and a plurality of first deep reinforcement learning models.
[0041] Among them, the hyperparameter combination can be a combination composed of one or more hyperparameters ai (i is an integer greater than or equal to 1). For example, an initial hyperparameter combination composed of hyperparameters a1, a2, and a3, an initial hyperparameter combination composed of hyperparameters a2, a4, and ai (i≠2 and i≠4), etc. The hyperparameter combination may specifically include at least one hyperparameter, such as the learning rate.
[0042] In the embodiments of the present application, multiple different initial hyperparameter combinations (e.g., 8 groups) can be obtained through multiple random samplings in advance within the reasonable value range of hyperparameters related to machine learning. Multiple deep reinforcement learning (Deep RL) models can be obtained according to different application scenarios. In the embodiments of the present application, the deep reinforcement model can be referred to as the first deep reinforcement learning model.
[0043] Specifically, when a deep reinforcement learning model with high training performance and good effect is required, multiple initial hyperparameter combinations are first obtained. At the same time, multiple first deep reinforcement learning models are obtained to serve as the basis for subsequent training models.
[0044] Among them, the first deep reinforcement model can be a model based on Off-line Deep RL (offline deep reinforcement learning). It can be understood that the Off-line Deep RL method is mainly based on historical offline data, and an optimal strategy is found by optimizing in the historical data. Off-line Deep RL techniques can include BCQ (Batch-Constrained deep Q-learning), Bear (Bootstrapping error accumulation reduction), BRAC (Behavior Regularized Offline Reinforcement Learning), CQL (Conservative Q-Learning), and other techniques.
[0045] S102, Train multiple first deep reinforcement learning models using multiple hyperparameters in the initial hyperparameter combination to obtain training evaluation indicators corresponding to the multiple first deep reinforcement learning models respectively.
[0046] Among them, the training evaluation indicator can be an indicator used to characterize the training effect or performance level of the deep reinforcement learning model.
[0047] Specifically, after obtaining multiple initial hyperparameter combinations and multiple first deep reinforcement learning models, each set of initial hyperparameters in the initial hyperparameter combination can be used to train multiple first deep reinforcement learning models respectively to obtain each trained first deep reinforcement learning model. Then, the effect or performance of each trained first deep reinforcement learning model can be evaluated through a model evaluator to obtain a training evaluation indicator corresponding to each first deep reinforcement learning model.
[0048] Among them, the model evaluator can use a simulation simulator or other offline strategies to evaluate the model. For example, use offline data to model the environment to obtain a simulation environment, and then let the model interact with the simulation environment to obtain the performance of the model in the simulation environment as the model evaluation. Or, use an offline reinforcement learning evaluation algorithm to evaluate the model.
[0049] It should be noted that the method for obtaining the training evaluation indicators corresponding to multiple first deep reinforcement learning models in the embodiments of the present application can also be other methods in the related art, as long as step S102 can be implemented. The embodiments of the present application do not limit this.
[0050] S103. Select a second deep reinforcement learning model from multiple first deep reinforcement learning models according to the training evaluation indicators.
[0051] Specifically, after obtaining the training evaluation indicators corresponding to multiple first deep reinforcement learning models, training evaluation indicators that meet the requirements can be selected from each training evaluation indicator according to the requirements, and the first deep reinforcement learning model corresponding to the selected training evaluation indicator, that is, the training evaluation indicator that meets the requirements, is used as the second deep reinforcement learning model.
[0052] It should be noted that the value range of the training evaluation indicator can be greater than or equal to 0 and less than or equal to 1. The larger the training evaluation indicator, the higher the training performance of the model; the smaller the training evaluation indicator, the lower the training performance of the model.
[0053] For example, if the training evaluation indicators corresponding to the first deep reinforcement learning models: M1, M2, M3, and M4 are 0.7 (good effect or performance), 0.2 (poor performance), 0.9 (very good performance), and 0.3 (poor performance) respectively, then in order to improve the performance of the deep reinforcement learning model, models M1 and M3 can be selected as the second deep reinforcement learning models, or model M3 can be selected as the second deep reinforcement learning model. Which model is specifically selected can be determined according to specific requirements.
[0054] It should be noted that in step S103, there may be training evaluation indicators that meet the requirements. At this time, the second deep reinforcement learning model can be selected according to the training evaluation indicators; there may also be no training evaluation indicators that meet the requirements. To avoid this phenomenon, a larger number of initial hyperparameter combinations and multiple first deep reinforcement learning models can be obtained in step S101 above.
[0055] S104. Optimize the initial hyperparameter combination by using multiple target hyperparameters corresponding to the second deep reinforcement learning model to form a target hyperparameter combination.
[0056] In the embodiments of the present application, the hyperparameters of each second deep reinforcement learning model can be referred to as target hyperparameters.
[0057] It should be noted that since Deep RL is greatly affected by hyperparameters during training, and at the same time, the optimization amplitude of the model based on Off-line Deep RL for the existing policy is limited, hyperparameters often affect the effectiveness of the model (that is, whether it can achieve a better level than the existing policy). Therefore, it is extremely important to tune the parameters of the Off-line Deep RL model. However, due to the large number of hyperparameters involved in Off-line Deep RL and the phenomenon of mutual coupling between hyperparameters, tuning the parameters of the Off-line Deep RL model is a very difficult task. Therefore, a hyperparameter optimization method needs to be introduced.
[0058] It should be noted that the basic methods of hyperparameter optimization can be divided into two types: parallel search and sequential optimization. Parallel search is to set multiple hyperparameter combinations for training at the same time, such as grid search and random search. Among them, grid search means that after setting the value range of hyperparameters, all possible hyperparameter combination configurations are traversed; random search means that when the number of hyperparameters is relatively large, sample points are randomly selected within the value range of hyperparameters. Sequential optimization is to combine the previous training experience and set better hyperparameter combinations for the next training iteration, such as Bayesian optimization method, Bandit-Based Algorithm Selection Methods, etc. Both parallel search and sequential optimization have their own drawbacks. The disadvantage of parallel search is that it does not utilize the parameter optimization information between models; while the sequential optimization process consumes a large amount of computing time. In the Off-line Deep RL scenario, due to the large number of hyperparameters, both will cause too low training efficiency and affect the entire training time of the model.
[0059] In view of this, the present application optimizes the initial hyperparameter combination through step S104.
[0060] Specifically, after determining the second deep reinforcement learning model, multiple target hyperparameters corresponding to the second deep reinforcement learning model can be determined, and the initial hyperparameter combination is optimized using the multiple target hyperparameters. The initial hyperparameter combination after optimization is the target hyperparameter combination, thereby obtaining the target hyperparameter combination.
[0061] For example, if the target hyperparameter combinations corresponding to the second deep reinforcement learning models M1 and M3 are a1, a3 and a1, a4, a5 respectively, then the target hyperparameters a1, a3, a4 and a5 are used to optimize multiple initial hyperparameter combinations to obtain the target hyperparameter combination. For example, one of the initial hyperparameter combinations a1, a2 and a4 is optimized into the target hyperparameter combination a1, a2, a3, a4 and a5.
[0062] In step S104, optimizing the initial hyperparameter combination based on multiple target hyperparameters requires fewer hyperparameters compared to the optimization methods of parallel search and sequential optimization in the related art, shortening the training time and thus improving the training efficiency.
[0063] S105. Train the second deep reinforcement learning model using multiple hyperparameters among the target hyperparameter combinations to obtain the target deep reinforcement learning model.
[0064] Specifically, after determining the target hyperparameter combination, use multiple hyperparameters among the target hyperparameters to train the second deep reinforcement learning model selected in step S103 to obtain the target deep reinforcement learning model.
[0065] Specifically, the second deep reinforcement learning model can be iteratively trained until a target deep reinforcement learning model is obtained when it converges.
[0066] For example, if the target hyperparameter combination is {a1, a3, a4, a5} and the second deep reinforcement learning models are M1 and M3, then use a1, a3, a4 and a5 to iteratively train the second deep reinforcement learning models M1 and M3 until the models converge to obtain a target deep reinforcement learning model.
[0067] That is to say, by executing the above steps S101 to S105, based on multiple initial hyperparameter combinations and multiple first deep reinforcement learning models, the training evaluation indicators corresponding to the multiple first deep reinforcement learning models can be obtained first, and then the second deep reinforcement learning model is selected from the multiple first deep reinforcement learning models according to the training evaluation indicators. After optimizing the initial hyperparameter combination with the multiple target hyperparameters corresponding to the second deep reinforcement learning model, the second deep reinforcement learning model is iteratively trained until a target deep reinforcement learning model with high performance and good effect is obtained.
[0068] In the embodiment of the present application, such as Figure 2As shown, first, a suitable Off-line DeepRL can be obtained based on a specific scenario, and then the hyperparameters of each model are optimized. After the optimization is completed, the trained models can be evaluated using a model evaluator, and the model with the best performance can be selected for output, thereby obtaining the target deep reinforcement learning model.
[0069] In the training method of the deep reinforcement learning model based on hyperparameter optimization according to the embodiments of the present application, based on multiple initial hyperparameter combinations and multiple deep reinforcement learning models, the initial hyperparameter combinations are optimized to form target hyperparameter combinations, and then multiple hyperparameters among the target hyperparameter combinations are used to train the target deep reinforcement learning model. Thus, combining hyperparameter optimization with model training to implement the training of the deep reinforcement learning model can not only train a deep reinforcement learning model with higher performance, but also enable the trained model to adapt to a wider range of application scenarios.
[0070] When performing the optimization process of hyperparameters in the above step S104, in order to improve the effectiveness of the optimization process, in an embodiment of the present application, as Figure 3 shown, the above step S104 may include the following steps S301 to S303:
[0071] S301, determine the hyperparameter set to which the initial hyperparameter combination belongs.
[0072] Specifically, after determining the second deep reinforcement learning model, determine the hyperparameter set {a1, a2,..., an} to which the initial hyperparameter combination belongs.
[0073] S302, supplement and add multiple target hyperparameters to the hyperparameter set to obtain a target hyperparameter set.
[0074] It should be noted that among the multiple target hyperparameters corresponding to the second deep reinforcement learning model, there may be target hyperparameters belonging to the hyperparameter set, or there may be target hyperparameters not belonging to the hyperparameter set.
[0075] Specifically, after determining the hyperparameter set {a1, a2,..., an}, the target hyperparameters among the multiple target hyperparameters corresponding to the second deep reinforcement learning model that do not belong to the hyperparameter set are supplemented and added to the hyperparameter set to obtain a target hyperparameter set.
[0076] For example, if the multiple target hyperparameters corresponding to the second deep reinforcement learning model are a2, a3, a5, and a6, and the hyperparameter set is {a1, a2, a3, a4}, then a5 and a6 among the multiple target hyperparameters can be supplemented and added to the hyperparameter set to obtain the target hyperparameter set {a1, a2, a3, a4, a5, a6}.
[0077] S303. Select at least some hyperparameters from the set of target hyperparameters, and form a target hyperparameter combination according to the at least some hyperparameters.
[0078] Specifically, after determining the set of target hyperparameters {a1, a2, …, am} (m ≥ n), at least some hyperparameters can be selected from the target hyperparameters, and a target hyperparameter combination can be formed according to the at least some hyperparameters.
[0079] For example, if the set of target hyperparameters is {a3, a5, a6, a7, a8, a9}, that is, there are 6 target hyperparameters, then at least 2 target hyperparameters can be selected from the 6 target hyperparameters to form a target hyperparameter combination, such as the combination of a3, a5 and a6, the combination of a5 and a6, the combination of a7, a8 and a9, etc.
[0080] That is to say, the optimization and expansion of the hyperparameter set are realized through the target hyperparameters, and then at least one target hyperparameter combination is generated according to some hyperparameters in the optimized and expanded hyperparameter set. Thus, not only the hyperparameter optimization process in the model training is realized, but also the effectiveness of the optimization process is improved.
[0081] It should be noted that when performing the above step S105, the second deep reinforcement learning model can be iteratively trained until a target deep reinforcement learning model is obtained when it converges. In order to improve the efficiency of hyperparameter optimization and generate a target hyperparameter combination with better effects, the above step S104 can be re-executed after each training, that is, the initial hyperparameter combination is optimized using multiple target hyperparameters corresponding to the second deep reinforcement learning model, or the above step S104 can be re-executed when the number of training times reaches a certain number. Based on this, the embodiments of the present application propose the following embodiments:
[0082] In an embodiment of the present application, the training method of the deep reinforcement learning model based on hyperparameter optimization may further include: when the number of times of training the second deep reinforcement learning model reaches the set number of iterations, re-optimize the initial hyperparameter combination using multiple target hyperparameters corresponding to the second deep reinforcement learning model.
[0083] Among them, the set number of iterations can be set in advance, or determined based on the current performance of the target deep learning model, or can be determined by other means, and the embodiments of the present application do not limit this.
[0084] Specifically, when iteratively training the second deep reinforcement learning model, the number of iterative training times can be counted. When the number of iterative training times reaches the set number of iterative times, the initial hyperparameter combination is re-optimized using multiple target hyperparameters corresponding to the second deep reinforcement learning model to obtain the target hyperparameter combination. Then, step S105 can be continued, that is, the second deep reinforcement learning model is trained using multiple hyperparameters among the target hyperparameter combination to obtain the target deep reinforcement learning model.
[0085] Thus, during the iterative training of the deep reinforcement learning model, at least one hyperparameter optimization process can be performed, which not only ensures that the reinforcement learning model has enough update and iteration steps to reach a fully convergent state, but also continuously adjusts the hyperparameters during this process to "carefully select" the model, thereby improving the hyperparameter optimization efficiency and generating a target hyperparameter combination with better effects.
[0086] As Figure 4 shown, the above step S105 may include the following steps:
[0087] S401, iteratively train the second deep reinforcement learning model using multiple hyperparameters among the target hyperparameter combination to obtain the predicted value output by the second deep reinforcement learning model.
[0088] Specifically, after forming the target hyperparameter combination, the second deep reinforcement learning model is iteratively trained using multiple hyperparameters among the target hyperparameter combination to obtain the predicted value output by the second deep reinforcement learning model.
[0089] S402, if the loss value between the predicted value and the calibration value meets the loss threshold, the trained second deep reinforcement learning model is used as the target deep reinforcement learning model.
[0090] Among them, the calibration value can be the predicted value of the model output that is used to characterize the better performance of the second deep reinforcement learning model. The loss threshold can refer to the maximum loss value between the predicted value and the calibration value, that is, if the loss value between the predicted value and the calibration value is the loss threshold, the performance of the second deep reinforcement learning model is better.
[0091] Specifically, when iteratively training the second deep reinforcement learning model to obtain the predicted value output by the second deep reinforcement learning model, the loss value between the predicted value and the calibration value can be calculated after each training. If the loss value meets the loss threshold, that is, the loss value is less than or equal to the loss threshold, the second deep reinforcement learning model obtained from the current training is used as the target deep reinforcement learning model; if the loss value does not meet the loss threshold, that is, the loss value is greater than the loss threshold, iterative training can be continued, or the initial hyperparameter combination can be optimized using multiple target hyperparameters corresponding to the current second deep reinforcement learning model to obtain the target hyperparameter combination. Then, steps S401 and S402 are executed above until the loss value meets the loss threshold.
[0092] For example, assume the loss threshold is 0.2. If after 3 trainings, the loss value between the predicted value output by the second deep reinforcement learning model and the calibration value is less than or equal to 0.2, the second deep reinforcement learning model obtained from the 3rd training is used as the target deep reinforcement learning model.
[0093] In this embodiment, whenever a certain number of steps of iterative training of the second deep reinforcement learning model is completed, a checkpoint can be set. By comparing with the performance of other models, the current model parameters are adjusted. If the current model performs well, continue training until the current model performs very well; if not, replace it with a better parameter configuration in other models and add random perturbations on this basis to continue training.
[0094] Thus, screening the target deep reinforcement learning model according to the predicted value output by the second deep reinforcement learning model improves the reliability of model screening.
[0095] It should be noted that the number of second deep reinforcement learning models screened according to the training evaluation index when executing step S103 above may be at least one. Therefore, the first deep reinforcement learning model with the training evaluation index ranked among the top in the multiple first deep reinforcement learning models can be used as the second deep reinforcement learning model.
[0096] That is, in an embodiment of the present application, when the number of second deep reinforcement learning models is the set number k, as Figure 5 shown, the above step S103 may include the following steps S501 to S502:
[0097] S501, rank the multiple first deep reinforcement learning models according to the training evaluation index.
[0098] It can be understood that the set number can be less than or equal to the number of first deep reinforcement learning models.
[0099] Specifically, after obtaining the training evaluation metrics corresponding to multiple first deep reinforcement learning models, sort the corresponding first deep reinforcement learning models according to the level of each training evaluation metric to obtain the sorted first deep reinforcement learning models. Specifically, the higher the training evaluation metric, the more forward the sorting of the corresponding second deep reinforcement learning model can be.
[0100] S502. Use the set number of first deep reinforcement learning models ranked at the front as the second deep reinforcement learning models.
[0101] Specifically, after sorting the first deep reinforcement learning models according to the training evaluation metrics, use the top k first deep reinforcement learning models as the second deep reinforcement learning models, thereby completing the screening of the second deep reinforcement learning models, and then execute steps S104 and S105.
[0102] For example, if the number of second deep reinforcement learning models to be screened is 2, and the training evaluation metrics of 4 first deep reinforcement learning models M1, M2, M3, and M4 are 0.1, 0.3, 0.8, and 0.7 respectively, then sort M1, M2, M3, and M4 as: M3 → M4 → M2 → M1. After that, the top 2 models, namely M3 and M4, can be used as the second deep reinforcement learning models.
[0103] Furthermore, in practical applications, the set number k in the embodiments of the present application is a parameter that can be adjusted according to needs. In an embodiment of the present application, the above step S103 may further include: determining the index performance requirements for training the first deep reinforcement learning models; adaptively adjusting the set number according to the index performance requirements.
[0104] Specifically, before performing the above step S502, the set number k can be determined according to the index performance requirements for training the first deep reinforcement learning models, and then k second deep reinforcement learning models are screened according to the training evaluation metrics. It should be noted that the index performance requirements are not fixed and may change with the index performance requirements. Therefore, it is necessary to determine the index performance requirements for training the first deep reinforcement learning models in real time, and then adaptively adjust the set number k according to the index performance requirements.
[0105] It should be noted that a larger set number k can increase the exploratory nature of hyperparameter optimization and avoid getting stuck in a local optimum situation, while a smaller set number k can accelerate convergence and thus speed up the screening speed.
[0106] Therefore, first determine the set number according to the performance index requirements, sort the first deep reinforcement learning models based on the training evaluation index, and directly use the first deep reinforcement learning models ranked among the set number as the second deep reinforcement learning models, thereby improving the reliability and convenience of model screening.
[0107] It should be noted that considering that the pure series optimization process takes too long, and at the same time, pure parallel search requires a large amount of training resources (such as GPU (Graphics Processing Unit), memory, CPU (Central Processing Unit), etc.), the embodiment of the present application introduces a method of coordinating parallel search and sequence optimization to ensure that the purpose of hyperparameter optimization is achieved with as short a computing time as possible under limited training resources.
[0108] That is, in an embodiment of the present application, training multiple first deep reinforcement learning models in the above step S102 may include: training multiple first deep reinforcement learning models in a parallel training manner.
[0109] Specifically, when training the first reinforcement learning model based on each set of hyperparameters in the initial hyperparameter combination, multiple GPUs are introduced for parallel training to improve the training speed and generate a model combination.
[0110] Based on the above embodiments, the following combines Figure 6 Describe a training method for a deep reinforcement learning model based on hyperparameter optimization in an example of the present application:
[0111] As Figure 6 shown, the main process of this method can be:
[0112] The first step is to give a historical data set for training an offline reinforcement learning model and initialize the hyperparameter combination.
[0113] The second step is to train a model evaluator based on the offline data.
[0114] Solution 1: Use the offline data to model the environment to obtain a simulation environment. Then, the model and the simulation environment are interacted to obtain the performance of the model in the simulation environment as the model evaluation.
[0115] Solution 2: Use an offline reinforcement learning evaluation algorithm to evaluate the model.
[0116] The third step is to train the reinforcement learning model (i.e., the first deep reinforcement learning model) based on each set of hyperparameters in the initial hyperparameter combination, and introduce multiple GPUs for parallel training to improve the speed and generate a model combination.
[0117] In the fourth step, use the model evaluator to evaluate the models trained with multiple groups of initial hyperparameters, and retain the superior models with good effects (i.e., the second deep reinforcement learning model).
[0118] In the fifth step, extract the hyperparameters of the superior model and perform hyperparameter expansion to generate new hyperparameter combinations (i.e., the target hyperparameter combinations).
[0119] In the sixth step, repeat the third step to the fifth step until a certain number of training steps.
[0120] In consideration of the characteristics of long training time and strong randomness in the reinforcement learning training process, the embodiment of this application introduces the Population Based Training (PBT) parameter optimization method. This method combines model training and hyperparameter optimization. It not only ensures that the reinforcement learning model has enough update and iteration steps to reach a fully convergent state, but also continuously adjusts the hyperparameters during this process to "carefully select" the model. The training method of the deep reinforcement learning model based on hyperparameter optimization provided by the embodiment of this application includes four processes: model preparation, hyperparameter optimization, model screening, and model output, achieving the effect of training a better model; moreover, based on the PBT algorithm and multiple GPUs, the hyperparameters of the reinforcement learning model are optimized during the training process to achieve the purpose of improving the training effect of offline reinforcement learning.
[0121] The following combines Figure 7 to describe a specific application scenario of the embodiment of this application:
[0122] It should be noted that the embodiment of this application can be applied to the combustion optimization control method for the industrial field, such as the thermal power generation optimization system. The state space of thermal power generating units is huge, and many internal physical and chemical reaction mechanisms are not yet clear. External inputs will be interfered by the environment, and different units have different characteristics. It is impossible for operators to fully understand their operating principles. Thermal power plants are equipped with different types of sensors everywhere, which can collect and summarize the readings in real time to the DCS system. The sensor sampling frequency is intensive, generating a large amount of data. Use big data to mine the internal laws, construct an offline reinforcement learning model, and define the optimization process as a multi-objective, multi-constraint, and high-dynamic decision-making process. By adjusting the control variables, the artificially defined objective function is optimized. The most direct goal is to improve the boiler efficiency and reduce pollutant emissions, generate more electricity with less coal, and produce less pollution.
[0123] In such scenarios, often a small improvement will bring considerable economic benefits, and the effect that the model can achieve is often affected by the hyperparameter settings. At this time, introducing the hyperparameter optimization method can help train a model with better optimization effect and improve the optimization effect. Through feature engineering, select several features from all the state features in the historical data to form a set s, such asFigure 7 As shown, the state feature s mainly includes data collected by sensors during the boiler combustion process, such as "furnace pressure", "furnace oxygen content", "steam flow rate", "steam temperature", "flue gas temperature", "flue gas pressure", "Nox content", "water wall temperature", "secondary air box pressure", "fan current", etc. The action a is defined as controllable variables that can be adjusted during the combustion process, such as "axial flow fan blade regulating valve", "flue gas damper", "primary and secondary air regulating valves", "desuperheating water regulating valve", etc. The historical operation records of the model in the past t time steps are {(s0, a0), (s1, a1)...(s t-1 , a t-1 )}. The information containing the historical operation records can be stored in the off-line data set D, and then a control model based on off-line Deep RL is constructed. The input of the off-line reinforcement learning model is the state feature s, and the output is the action a. Then, using the off-line data set, the model is trained by the training method proposed in the embodiments of the present application to output a higher-quality target deep reinforcement learning model.
[0124] The embodiments of the present application also propose a training device for a deep reinforcement learning model based on hyperparameter optimization, Figure 8 which is a schematic structural diagram of a training device for a deep reinforcement learning model based on hyperparameter optimization provided by the embodiments of the present application.
[0125] As Figure 8 shown, the training device 800 for a deep reinforcement learning model based on hyperparameter optimization includes: a first acquisition module 810, a first training module 820, a first screening module 830, a first processing module 840, and a second training module 850.
[0126] Among them, the first acquisition module 810 is used to acquire a plurality of initial hyperparameter combinations and a plurality of first deep reinforcement learning models; the first training module 820 is used to train the plurality of first deep reinforcement learning models with a plurality of hyperparameters in the initial hyperparameter combinations to obtain training evaluation indicators corresponding to the plurality of first deep reinforcement learning models respectively; the first screening module 830 is used to screen out a second deep reinforcement learning model from the plurality of first deep reinforcement learning models according to the training evaluation indicators; the first processing module 840 is used to optimize the initial hyperparameter combinations with a plurality of target hyperparameters corresponding to the second deep reinforcement learning model to form target hyperparameter combinations; and the second training module 850 is used to train the second deep reinforcement learning model with a plurality of hyperparameters in the target hyperparameter combinations to obtain a target deep reinforcement learning model.
[0127] In one embodiment of the present application, the first processing module 840 may include: a first determination unit configured to determine a hyperparameter set to which an initial hyperparameter combination belongs; a first addition unit configured to supplement and add a plurality of target hyperparameters to the hyperparameter set to obtain a target hyperparameter set; and a first selection unit configured to select at least some hyperparameters from the target hyperparameter set and form a target hyperparameter combination based on the at least some hyperparameters.
[0128] In one embodiment of the present application, the training apparatus 800 for a deep reinforcement learning model based on hyperparameter optimization may further include: a second processing module configured to, when the number of times of training the second deep reinforcement learning model reaches a set number of iterations, re-optimize the initial hyperparameter combination by using a plurality of target hyperparameters corresponding to the second deep reinforcement learning model.
[0129] In one embodiment of the present application, the second training module may include: a first training unit configured to iteratively train a second deep reinforcement learning model by using a plurality of hyperparameters in a target hyperparameter combination to obtain a predicted value output by the second deep reinforcement learning model; a first determination unit configured to, if a loss value between the predicted value and a calibration value satisfies a loss threshold, use the trained second deep reinforcement learning model as a target deep reinforcement learning model.
[0130] In one embodiment of the present application, the first training module may include: a second training unit configured to train a first deep reinforcement learning model corresponding to each hyperparameter combination in a parallel training manner.
[0131] In one embodiment of the present application, the second training module may further include: a third determination unit configured to determine an index performance requirement for training the first deep reinforcement learning model; a first adjustment unit configured to adaptively adjust a set number according to the index performance requirement.
[0132] It should be noted that other specific embodiments of the training apparatus for a deep reinforcement learning model based on hyperparameter optimization in the embodiments of the present application may refer to the specific embodiments of the training method for a deep reinforcement learning model based on hyperparameter optimization described above. To avoid redundancy, they will not be elaborated here.
[0133] The training apparatus for a deep reinforcement learning model based on hyperparameter optimization in the embodiments of the present application combines hyperparameter optimization and model training to implement the training of a deep reinforcement learning model. It can not only train a deep reinforcement learning model with better performance, but also enable the trained model to adapt to a wider range of application scenarios.
[0134] According to an embodiment of the present application, the present application further provides an electronic device, a readable storage medium, and a computer program product for a training method of a deep reinforcement learning model based on hyperparameter optimization. The following combines Figure 9A description will be given.
[0135] Figure 9 It is a block diagram of an electronic device for a method of training a deep reinforcement learning model based on hyperparameter optimization according to an embodiment of the present application.
[0136] As Figure 9 shown, the electronic device 900 may include: a memory 910 and at least one processor 920, and a bus 930 connecting different components (including the memory 910 and the processor 920).
[0137] A computer program is stored on the memory 910, and when the processor 920 executes the program, the method of training a deep reinforcement learning model based on hyperparameter optimization according to an embodiment of the present application is implemented.
[0138] The bus 930 represents one or more of several types of bus architectures, including a memory bus or a memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of the various bus architectures. By way of example, these architectures include, but are not limited to, Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MAC) bus, Enhanced ISA bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnection (PCI) bus.
[0139] The electronic device 900 typically includes a variety of computer system readable media. These media can be any available media accessible by the electronic device 800, including volatile and non-volatile media, removable and non-removable media.
[0140] The memory 910 may include computer system readable media in the form of volatile memory, such as Random Access Memory (RAM) 940 and / or cache memory 950. The electronic device 900 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, a storage system 960 may be used for reading and writing non-removable, non-volatile magnetic media ( Figure 9 not shown, commonly referred to as a "hard disk drive"). Although Figure 9Not shown in the figure, a disk drive for reading and writing to a removable non-volatile disk (such as a "floppy disk") and an optical disk drive for reading and writing to a removable non-volatile optical disk (such as: Compact Disc Read Only Memory (CD-ROM), Digital Video Disc Read Only Memory (DVD-ROM) or other optical media) can be provided. In these cases, each drive can be connected to the bus 930 through one or more data medium interfaces. The memory 910 may include at least one program product having a set (such as at least one) of program modules configured to perform the functions of the embodiments of the present invention.
[0141] A program / utility 980 having a set (at least one) of program modules 970 can be stored, for example, in the memory 910. Such program modules 970 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment. The program modules 870 generally perform the functions and / or methods in the embodiments described in the present invention.
[0142] The electronic device 900 can also communicate with one or more external devices 990 (such as a keyboard, a pointing device, a display 991, etc.), and can also communicate with one or more devices that enable a user to interact with the electronic device 900, and / or communicate with any device that enables the electronic device 900 to communicate with one or more other computing devices (such as a network card, a modem, etc.). Such communication can be carried out through the input / output (I / O) interface 992. Moreover, the electronic device 900 can also communicate with one or more networks (such as a Local Area Network (LAN), a Wide Area Network (WAN), and / or a public network, such as the Internet) through the network adapter 993. As shown in the figure, the network adapter 993 communicates with other modules of the electronic device 900 through the bus 930. It should be understood that although not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 900, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0143] The processor 920 executes various functional applications and data processing by running the programs stored in the memory 910, such as implementing the methods mentioned in the foregoing embodiments.
[0144] In the description of this specification, the descriptions with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms are not necessarily directed to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples.
[0145] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the present invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise specifically defined.
[0146] Any process or method description, whether in a flowchart or otherwise described herein, can be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing a customized logical function or process, and the scope of the preferred embodiments of the present invention includes additional implementations, where the functions may be performed in a substantially simultaneous manner or in an order opposite to that shown or discussed, according to the functions involved, which should be understood by those skilled in the art to which the embodiments of the present invention pertain.
[0147] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definable sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in conjunction with these instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection portion with one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other suitable processing as necessary, and then storing it in a computer memory.
[0148] It should be understood that various parts of the present invention can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having suitable combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0149] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the method of implementing the above embodiments can be completed by a program instructing relevant hardware, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.
[0150] In addition, in each embodiment of the present invention, each functional unit can be integrated into a processing module, or each unit can exist physically alone, or two or more units can be integrated into one module. The above-mentioned integrated module can be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.
[0151] The above-mentioned storage medium can be a read-only memory, a magnetic disk, an optical disk, etc. Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.
Claims
1. A training method for a deep reinforcement learning model based on hyperparameter optimization, the method being used for combustion optimization control in the industrial field, characterized in that, The method includes: obtaining a plurality of initial hyperparameter combinations and a plurality of first deep reinforcement learning models; training the plurality of first deep reinforcement learning models with a plurality of hyperparameters in the initial hyperparameter combinations to obtain training evaluation metrics respectively corresponding to the plurality of first deep reinforcement learning models, wherein the inputs of the plurality of first deep reinforcement learning models are state features, the outputs are actions, the state features include data collected by sensors during the boiler combustion process, and the actions include control variables that can be adjusted during the combustion process; screening out a second deep reinforcement learning model from the plurality of first deep reinforcement learning models according to the training evaluation metrics; optimizing the initial hyperparameter combinations with a plurality of target hyperparameters corresponding to the second deep reinforcement learning model to form target hyperparameter combinations; and training the second deep reinforcement learning model with a plurality of hyperparameters in the target hyperparameter combinations to obtain a target deep reinforcement learning model; wherein, the training of the plurality of first deep reinforcement learning models includes: training a first reinforcement learning model respectively based on each group of hyperparameters in the initial hyperparameter combinations, and introducing multiple GPUs for parallel training to generate a model combination.
2. The method according to claim 1, characterized in that, The optimizing the initial hyperparameter combinations with a plurality of target hyperparameters corresponding to the second deep reinforcement learning model to form target hyperparameter combinations includes: determining the hyperparameter set to which the initial hyperparameter combinations belong; supplementing and adding the plurality of target hyperparameters to the hyperparameter set to obtain a target hyperparameter set; and selecting at least part of the hyperparameters from the target hyperparameter set and forming the target hyperparameter combinations according to the at least part of the hyperparameters.
3. The method according to claim 1, characterized in that, It further includes: when the number of times of training the second deep reinforcement learning model reaches a set number of iterations, re-optimizing the initial hyperparameter combinations with a plurality of target hyperparameters corresponding to the second deep reinforcement learning model.
4. The method according to claim 1, characterized in that, The training the second deep reinforcement learning model with a plurality of hyperparameters in the target hyperparameter combinations to obtain a target deep reinforcement learning model includes: iteratively training the second deep reinforcement learning model with a plurality of hyperparameters in the target hyperparameter combinations to obtain predicted values output by the second deep reinforcement learning model; if the loss value between the predicted value and the calibration value meets a loss threshold, using the trained second deep reinforcement learning model as the target deep reinforcement learning model.
5. The method according to claim 1, characterized in that, The number of the second deep reinforcement learning models is a set number, and the screening out a second deep reinforcement learning model from the plurality of first deep reinforcement learning models according to the training evaluation metrics includes: sorting the plurality of first deep reinforcement learning models according to the training evaluation metrics; using the set number of first deep reinforcement learning models ranked at the front as the second deep reinforcement learning models.
6. The method according to claim 5, characterized in that, It further includes: determining the index performance requirements for training the first deep reinforcement learning model; adaptively adjusting the set number according to the index performance requirements.
7. A training device for a deep reinforcement learning model based on hyperparameter optimization, the device being used for combustion optimization control in the industrial field, characterized in that, The device includes: The first acquisition module is used to acquire multiple initial hyperparameter combinations and multiple first deep reinforcement learning models, where the inputs of the multiple first deep reinforcement learning models are state features, the outputs are actions, the state features include data collected by sensors during the boiler combustion process, and the actions include controllable variables that can be adjusted during the combustion process; The first training module is used to train the multiple first deep reinforcement learning models with multiple hyperparameters in the initial hyperparameter combination to obtain training evaluation metrics corresponding to the multiple first deep reinforcement learning models respectively; The first screening module is used to screen out a second deep reinforcement learning model from the multiple first deep reinforcement learning models according to the training evaluation metrics; The first processing module is used to optimize the initial hyperparameter combination with multiple target hyperparameters corresponding to the second deep reinforcement learning model to form a target hyperparameter combination; and The second training module is used to train the second deep reinforcement learning model with multiple hyperparameters in the target hyperparameter combination to obtain a target deep reinforcement learning model; Among them, the first training module includes: The second training unit is used to train the first reinforcement learning model respectively based on each group of hyperparameters in the initial hyperparameter combination, and introduce multiple GPUs for parallel training to generate a model combination.
8. The device according to claim 7, wherein The first processing module includes: The first determination unit is used to determine the hyperparameter set to which the initial hyperparameter combination belongs; The first addition unit is used to supplement and add the multiple target hyperparameters to the hyperparameter set to obtain a target hyperparameter set; and The first selection unit is used to select at least part of the hyperparameters from the target hyperparameter set and form the target hyperparameter combination according to the at least part of the hyperparameters.
9. The device according to claim 7, wherein It further includes: The second processing module is used to, when the number of times of training the second deep reinforcement learning model reaches the set number of iterations, re-optimize the initial hyperparameter combination with multiple target hyperparameters corresponding to the second deep reinforcement learning model.
10. The device according to claim 7, wherein The second training module includes: The first training unit is used to iteratively train the second deep reinforcement learning model with multiple hyperparameters in the target hyperparameter combination to obtain the predicted value output by the second deep reinforcement learning model; The first determination unit is used to, if the loss value between the predicted value and the calibration value meets the loss threshold, use the trained second deep reinforcement learning model as the target deep reinforcement learning model.
11. The device according to claim 7, wherein The number of the second deep reinforcement learning models is the set number, and the second training module includes: The first sorting unit is used to sort the multiple first deep reinforcement learning models according to the training evaluation metrics; The second determination unit is used to use the set number of first deep reinforcement learning models ranked in the front as the second deep reinforcement learning models.
12. The device according to claim 11, wherein The second training module further includes: The third determination unit is used to determine the index performance requirements for training the first deep reinforcement learning model; A first adjustment unit for adaptively adjusting the set number according to the performance requirements of the metrics.
13. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the training method of the deep reinforcement learning model based on hyperparameter optimization according to any one of claims 1-6.
14. A non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the training method of the hyperparameter-optimized deep reinforcement learning model according to any one of claims 1-6.
15. A computer program product, comprising a computer program, wherein When the computer program is executed by a processor, it implements the training method of the deep reinforcement learning model based on hyperparameter optimization according to any one of claims 1-6.
Citation Information
Patent Citations
An optimization method of vibration spectrum analysis model based on workflow
CN109299501A
Construction method and device of deep learning model for positioning multi-target feature points
CN110188780A