Method, apparatus, device, and storage medium for determining hyperparameters of a deep learning model
Through distributed cluster node training and evaluation of hyperparameter groups, the deep learning model hyperparameters are automatically configured, which solves the problem of inefficient sorting of product search results, and realizes personalized and efficient model training.
Patent Information
- Application Number
- CN202210051156.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-17
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-01-17
AI Technical Summary
The lack of effective deep learning model hyperparameter determination method in the prior art has led to inefficient sorting of product search results and cannot meet personalized needs.
Through the distributed cluster node training and evaluation of multiple hyperparameter groups, storing and optimizing the evaluation value, the optimal training model and its corresponding hyperparameter group are finally output to achieve automated hyperparameter configuration.
It improves the training efficiency and accuracy of deep learning models, meets the personalized needs of different regions, and reduces the time and labor cost of manually configuring hyperparameters.
Smart Images

Figure CN114418091B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of deep learning, and particularly relates to a method, apparatus, device, and storage medium for determining hyperparameters of a deep learning model. Background Art
[0002] With the rapid development of e-commerce, the number of similar products on e-commerce platforms is huge. Relying solely on product retrieval systems cannot meet the personalized needs of users. Therefore, deep learning models based on the ranking of product retrieval results are used in e-commerce platforms. Due to different consumption levels, shopping habits, and available products in different regions, e-commerce platforms need to customize different deep learning models for each region. A good deep learning model needs to balance prediction accuracy, real-time performance, and training duration, all of which are directly related to the configuration of hyperparameters.
[0003] To make each deep learning model perform excellently, it is necessary to configure the most suitable hyperparameters for each deep learning model. However, manually configuring hyperparameters for deep learning models will consume a large amount of time and manpower. Therefore, automatic hyperparameter optimization is proposed, aiming to configure hyperparameters without manual participation.
[0004] Currently, there is no method with better effects for determining hyperparameters of deep learning models for ranking product retrieval results. Summary of the Invention
[0005] This application provides a method, apparatus, device, and storage medium for determining hyperparameters of a deep learning model, aiming to solve the technical problem that there is currently no method with better effects for determining hyperparameters of deep learning models for ranking product retrieval results.
[0006] The first aspect of this application provides a method for determining hyperparameters of a deep learning model, including:
[0007] S101. Obtain model data to be trained, where the model data includes a target model, training data, and test data;
[0008] S102. Obtain a number of hyperparameter groups;
[0009] S103. Based on the hyperparameter groups, the training data, and the test data, train and evaluate the target model through a number of distributed cluster sub-nodes to obtain a number of training models and evaluation values of the hyperparameter groups;
[0010] S104. Store the hyperparameter groups and the evaluation values of the hyperparameter groups, and store the optimal training model from the number of training models;
[0011] S105. Optimize the stored hyperparameter groups, and select several new hyperparameter groups as the hyperparameter groups for the next round of training;
[0012] S106. Repeat steps S102 to S105. When the number of training times or the training duration reaches a preset value, output the optimal training model and its corresponding hyperparameter group.
[0013] In this application, several hyperparameter groups are first obtained, different hyperparameter groups are configured for the target model, and then the target models configured with different hyperparameter groups are respectively trained and evaluated through training data such as product retrieval results and test data. After each round of training, the hyperparameter groups, the evaluation values of the hyperparameter groups, and the optimal training model in this round of training task are stored. Then, the stored hyperparameter groups are optimized to select the hyperparameter groups required for the next round of training task. By continuously training in this way, better hyperparameter groups are obtained. When the number of training times or the training duration reaches a preset value, the optimal training model and its corresponding hyperparameter group are output to determine the hyperparameter group, thereby solving the technical problem that there is currently no method for determining hyperparameters of a deep learning model with better performance for sorting product retrieval results.
[0014] Optionally, the optimizing the stored hyperparameter groups and selecting new hyperparameter groups as the hyperparameter groups for the next round of training includes:
[0015] Select two hyperparameter groups with relatively better evaluation values;
[0016] Divide a positive sampling region in the sampling space;
[0017] Sample in the positive sampling region with a first preset probability and sample in the global sampling region with a second preset probability to select new hyperparameter groups as the hyperparameter groups for the next round of training.
[0018] Optionally, the dividing a positive sampling region in the sampling space includes:
[0019] Classify the two hyperparameter groups to obtain numerical type hyperparameters and categorical type hyperparameters;
[0020] Draw an axis-parallel box with the numerical type hyperparameters as the edges, and the region covered by the axis-parallel box is the positive sampling region of the numerical type hyperparameters;
[0021] Use the categorical type hyperparameters as the positive sampling region of the categorical type hyperparameters.
[0022] Optionally, the first preset probability is greater than the second preset probability.
[0023] Optionally, the obtaining several hyperparameter groups includes:
[0024] Select a number of hyperparameter groups for the first evaluation according to storage experience.
[0025] Optionally, training and evaluating the target model through a number of distributed cluster sub-nodes based on the hyperparameter groups, the training data, and the test data to obtain evaluation values of a number of trained models and hyperparameter groups, including:
[0026] Allocate the target model, training data, and test data to each of the distributed cluster sub-nodes;
[0027] Configure one of the hyperparameter groups for each of the distributed cluster sub-nodes;
[0028] Train and evaluate the target model through a number of the distributed cluster sub-nodes to obtain evaluation values of a number of trained models and hyperparameter groups.
[0029] The second aspect of the present application provides an apparatus for determining hyperparameters of a deep learning model, including:
[0030] A first acquisition unit for acquiring model data to be trained, where the model data includes a target model, training data, and test data;
[0031] A second acquisition unit for acquiring a number of hyperparameter groups;
[0032] A training unit for training and evaluating the target model through a number of distributed cluster sub-nodes based on the hyperparameter groups, the training data, and the test data to obtain evaluation values of a number of trained models and hyperparameter groups;
[0033] A storage unit for storing the hyperparameter groups and the evaluation values of the hyperparameter groups, and storing the optimal trained model from a number of the trained models;
[0034] An optimization unit for optimizing the stored hyperparameter groups and selecting a number of new hyperparameter groups as the hyperparameter groups for the next round of training;
[0035] An output unit for outputting the optimal trained model and its corresponding hyperparameter group when the number of training times or the training duration reaches a preset value.
[0036] The third aspect of the present application provides an electronic device, including a processor and a memory storing a computer program, and when the processor executes the computer program, the steps of the method for determining hyperparameters of a deep learning model as described in the first aspect are implemented.
[0037] The fourth aspect of the present application provides a computer-readable storage medium, on which a computer program is stored, characterized in that when the computer program is executed by a processor, the steps of the method for determining hyperparameters of the deep learning model as described in the first aspect are implemented. Description of the Drawings
[0038] Figure 1 It is a schematic flowchart of a method for determining hyperparameters of a deep learning model provided by an embodiment of the present application;
[0039] Figure 2 It is a schematic diagram of the division of positive sampling regions provided by an embodiment of the present application;
[0040] Figure 3 It is a training schematic diagram of multiple distributed cluster sub-nodes provided by an embodiment of the present application;
[0041] Figure 4 It is a schematic structural diagram of a device for determining hyperparameters of a deep learning model provided by an embodiment of the present application;
[0042] Figure 5 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed Embodiments
[0043] The embodiments of the present application provide a method, device, equipment, and storage medium for determining hyperparameters of a deep learning model, which are used to solve the technical problem that there is currently no relatively optimal method for determining hyperparameters of a deep learning model for ranking commodity retrieval results.
[0044] In order to make the objectives, features, and advantages of the present application more obvious and understandable, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the embodiments described below are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0045] Please refer to Figures 1 to 3 , the embodiments of the present application provide a method for determining hyperparameters of a deep learning model, including:
[0046] Step S101, obtain model data to be trained, where the model data includes a target model, training data, and test data.
[0047] When performing deep learning, first, the model data to be trained needs to be obtained. The model data to be trained generally includes a target model to be trained, training data such as commodity retrieval result data, and test data for verifying the model.
[0048] Step S102: Obtain a number of hyperparameter groups.
[0049] Before training the target model, a hyperparameter group needs to be configured for each target model. Therefore, before training, a number of hyperparameter groups need to be obtained first.
[0050] Step S103: Based on the hyperparameter groups, training data, and test data, train and evaluate the target model through a number of distributed cluster sub-nodes to obtain a number of trained models and evaluation values of the hyperparameter groups.
[0051] For each round of training, the target model needs to be trained through a number of distributed cluster sub-nodes based on the hyperparameter group, training data, and test data of that round. The hyperparameter groups for different training tasks are different. Therefore, when entering the next round of training, a new hyperparameter group needs to be replaced.
[0052] Step S104: Store the hyperparameter groups and the evaluation values of the hyperparameter groups, and store the optimal trained model from a number of trained models.
[0053] It should be noted that after each round of training, the hyperparameter groups, the evaluation values of the hyperparameter groups, and the model type (i.e., the optimal trained model) need to be stored to provide data support for optimizing the hyperparameter groups in the future.
[0054] Step S105: Optimize the stored hyperparameter groups and select new hyperparameter groups as the hyperparameter groups for the next round of training.
[0055] The hyperparameter groups configured for training the model in this embodiment are all optimized and are optimized based on the hyperparameter groups of the previous round. Therefore, the determined hyperparameter groups have better effects.
[0056] Step S106: Repeat steps S102 to S105. When the number of training times or the training duration reaches the preset value, output the optimal trained model and its corresponding hyperparameter group.
[0057] Repeat steps S102 to S105, continuously train and evaluate the model and optimize the hyperparameter groups. When the number of training times or the training duration meets the preset value, output the optimal trained model and its corresponding hyperparameter group. It can be understood that both the number of training times and the training duration can be set according to the actual training situation of the model. Since the optimal trained model is stored after each round of training, that is, if the current model is better than the stored model, the stored model is replaced; otherwise, the stored model remains unchanged. Therefore, at the end of training, the optimal model in multiple rounds of training is stored in the storage unit to output the optimal model and its corresponding hyperparameter group.
[0058] Furthermore, step S105 includes:
[0059] Select two sets of hyperparameter groups with relatively better evaluation values.
[0060] Divide the positive sampling area in the sampling space.
[0061] Sample in the positive sampling area with the first preset probability and sample in the global sampling area with the second preset probability to select a new set of hyperparameters as the hyperparameters for the next round of training.
[0062] It should be noted that the first preset probability is greater than the second preset probability, and the sum of the first preset probability and the second preset probability is 1. For example, sample in the positive sampling area with probability a and sample in the global sampling area with probability (1 - a). Different sampling probabilities are applied to different sampling areas, and mainly the positive sampling area is given priority, making the sampling accuracy higher.
[0063] Furthermore, dividing the positive sampling area in the sampling space includes:
[0064] Classify the two sets of hyperparameter groups to obtain numerical type hyperparameters and categorical type hyperparameters.
[0065] Use the numerical type hyperparameters to draw axis-parallel boxes at the edges, and the area covered by the axis-parallel boxes is the positive sampling area of the numerical type hyperparameters.
[0066] Take the categorical type hyperparameters as the positive sampling area of the categorical type hyperparameters.
[0067] It can be understood that different methods are used to divide the positive sampling area for different types of hyperparameters, making the targeting stronger when selecting the hyperparameter groups for the next round. The two sets of hyperparameter groups with relatively better evaluation values are {a4, b2, c3, d4} and {a2, b4, c3, d5} respectively. Hyperparameter A and hyperparameter B are numerical type hyperparameters, and hyperparameter C and hyperparameter D are categorical type hyperparameters. The positive sampling area is Figure 2 The dotted box shown.
[0068] Furthermore, step S102 includes:
[0069] Select several hyperparameter groups for the first evaluation according to the stored experience.
[0070] It should be noted that during the first evaluation, that is, in the first round of training, the selection method of the hyperparameter groups can be based on the stored experience. That is, judge whether the stored experience is empty. If not, select several hyperparameter groups for the first evaluation according to the experience. If so, randomly initialize several hyperparameter groups. The hyperparameter groups selected according to the experience are usually close to the model data to be trained.
[0071] Furthermore, step S103 includes:
[0072] Allocate a target model, training data, and test data to each distributed cluster sub-node;
[0073] Configure a hyperparameter group for each distributed cluster sub-node;
[0074] Train and evaluate the target model through several distributed cluster sub-nodes to obtain several trained models and evaluation values of the hyperparameter groups.
[0075] As Figure 3 shown, multiple distributed cluster sub-nodes train the target model based on different hyperparameter groups respectively. The more distributed cluster sub-nodes there are, the faster the synchronous training speed is.
[0076] In this embodiment, several hyperparameter groups are randomly obtained first, different hyperparameter groups are configured for the target model, and then the target models configured with different hyperparameter groups are trained by training data such as commodity retrieval results and test data respectively. After each round of training, the hyperparameter group, the evaluation value of the hyperparameter group, and the optimal trained model in this round of training are stored, and then the stored hyperparameter group is optimized to select the hyperparameter group required for the next round of training. Through continuous training like this, a hyperparameter group with better effects is obtained. When the number of training times or the training duration reaches the preset value, the optimal trained model and its corresponding hyperparameter group are output to determine the hyperparameter group, thus solving the technical problem that there is currently no method for determining hyperparameters of a deep learning model for sorting commodity retrieval results with better effects.
[0077] The above is a detailed description of an embodiment of a method for determining hyperparameters of a deep learning model provided by this application. The following is a detailed description of an embodiment of a device for determining hyperparameters of a deep learning model provided by this application. The device for determining hyperparameters of a deep learning model described below can be correspondingly referred to with the method for determining hyperparameters of a deep learning model described above.
[0078] Please refer to Figure 4 , this application embodiment provides a device for determining hyperparameters of a deep learning model, including:
[0079] A first acquisition unit 401, configured to acquire model data to be trained, where the model data includes a target model, training data, and test data.
[0080] A second acquisition unit 402, configured to acquire several hyperparameter groups.
[0081] A training unit 403, configured to train and evaluate the target model through several distributed cluster sub-nodes based on the hyperparameter groups, training data, and test data to obtain several trained models and evaluation values of the hyperparameter groups.
[0082] The storage unit 404 stores the hyperparameter group and the evaluation value of the hyperparameter group, and stores the optimal trained model from several trained models.
[0083] The optimization unit 405 is used to optimize the stored hyperparameter group and select several new hyperparameter groups as the hyperparameter groups for the next round of training.
[0084] The output unit 406 is used to output the optimal trained model and its corresponding hyperparameter group when the number of training times or the training duration reaches a preset value.
[0085] Further, the optimization unit 405 includes:
[0086] The first selection subunit is used to select two hyperparameter groups with better evaluation values.
[0087] The division subunit is used to divide a positive sampling region in the sampling space.
[0088] The sampling subunit is used to sample in the positive sampling region with a first preset probability and sample in the global sampling region with a second preset probability to select a new hyperparameter group as the hyperparameter group for the next round of training.
[0089] Further, the division subunit includes:
[0090] The classification subunit is used to classify the two hyperparameter groups to obtain numerical type hyperparameters and categorical type hyperparameters.
[0091] The first division subunit is used to outline an axis-parallel box with the numerical type hyperparameters as the edges, and the region covered by the axis-parallel box is the positive sampling region of the numerical type hyperparameters.
[0092] The second division subunit is used to use the categorical type hyperparameters as the positive sampling region of the categorical type hyperparameters.
[0093] Further, the first preset probability is greater than the second preset probability.
[0094] Further, the second acquisition unit 402 is specifically used to select several hyperparameter groups for the first evaluation according to the stored experience.
[0095] Further, the training unit 403 includes:
[0096] The allocation subunit is used to allocate the target model, training data, and test data to each of the distributed cluster sub-nodes.
[0097] The configuration subunit is used to configure one hyperparameter group for each of the distributed cluster sub-nodes.
[0098] A training and evaluation subunit, configured to train and evaluate the target model through a plurality of the distributed cluster sub-nodes, and obtain evaluation values of a plurality of trained models and hyperparameter groups.
[0099] In this embodiment, a plurality of hyperparameter groups are randomly obtained first, different hyperparameter groups are configured for the target model, and then the target models configured with different hyperparameter groups are respectively trained and evaluated through training data such as product retrieval results and test data. After each round of training, the hyperparameter group, the evaluation value of the hyperparameter group, and the optimal trained model in this round of training are stored, and then the stored hyperparameter group is optimized to select the hyperparameter group required for the next round of training. By continuously training in this way, a better hyperparameter group can be obtained. When the number of training times or the training duration reaches a preset value, the optimal trained model and its corresponding hyperparameter group are output to determine the hyperparameter group, thus solving the technical problem that there is currently no method for determining hyperparameters of a deep learning model with better performance for product retrieval result ranking.
[0100] Figure 5 An example of the physical structure diagram of an electronic device is shown as Figure 5 As shown, the present invention further provides an electronic device, which may include: a processor 310, a communication interface 320, a memory 330, and a communication bus 340. Among them, the processor 310, the communication interface 320, and the memory 330 complete communication with each other through the communication bus 340. The processor 310 can call the computer program in the memory 330 to execute the steps of a method for determining hyperparameters of a deep learning model, for example, including:
[0101] Obtain model data to be trained, where the model data includes a target model, training data, and test data;
[0102] Obtain a plurality of hyperparameter groups;
[0103] Through a plurality of distributed cluster sub-nodes, train and evaluate the target model based on the hyperparameter group, training data, and test data, and obtain evaluation values of a plurality of trained models and hyperparameter groups;
[0104] Store the hyperparameter group and the evaluation value of the hyperparameter group, and store the optimal trained model from a plurality of trained models;
[0105] Optimize the stored hyperparameter group, and select a plurality of new hyperparameter groups as the hyperparameter groups for the next round of training;
[0106] Repeat the above steps. When the number of training times or the training duration reaches a preset value, output the optimal trained model and its corresponding hyperparameter group.
[0107] In addition, when the logical instructions in the above-mentioned memory 330 can be implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The aforementioned storage medium includes: various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.
[0108] On the other hand, an embodiment of the present application also provides a computer-readable storage medium. The processor-readable storage medium stores a computer program, and the computer program is used to cause the processor to execute the steps of the methods provided in the above-mentioned various embodiments, for example, including:
[0109] Obtain model data to be trained, where the model data includes a target model, training data, and test data;
[0110] Obtain a number of hyperparameter groups;
[0111] Through a number of distributed cluster sub-nodes, train and evaluate the target model based on the hyperparameter groups, the training data, and the test data to obtain a number of training models and evaluation values of the hyperparameter groups;
[0112] Store the hyperparameter groups and the evaluation values of the hyperparameter groups, and store the optimal training model from a number of the training models;
[0113] Optimize the stored hyperparameter groups, and select a number of new hyperparameter groups as the hyperparameter groups for the next round of training;
[0114] Repeat the above steps. When the number of training times or the training duration reaches a preset value, output the optimal training model and its corresponding hyperparameter groups.
[0115] The processor-readable storage medium can be any available medium or data storage device that the processor can access, including but not limited to magnetic memories (such as floppy disks, hard disks, magnetic tapes, magneto-optical discs (MO), etc.), optical memories (such as CDs, DVDs, BDs, HVDs, etc.), and semiconductor memories (such as ROM, EPROM, EEPROM, non-volatile memories (NANDFLASH), solid-state drives (SSD)).
[0116] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative work.
[0117] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course also by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0118] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments or perform equivalent replacements for some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for determining hyperparameters of a deep learning model, characterized in that Including: S101. Obtain model data to be trained, where the model data includes a target model, training data, and test data, and the target model is a deep learning model for sorting commodity retrieval results; S102. Obtain a number of hyperparameter groups; S103. Based on the hyperparameter groups, the training data, and the test data, train and evaluate the target model through a number of distributed cluster sub-nodes to obtain a number of evaluation values of the training models and hyperparameter groups, where the number of the distributed cluster sub-nodes is positively correlated with the synchronous training speed; S104. Store the hyperparameter groups and the evaluation values of the hyperparameter groups, and store the optimal training model from a number of the training models, including: Select two groups of hyperparameter groups with relatively better evaluation values; Divide a positive sampling area in the sampling space, including: classifying the two groups of hyperparameter groups to obtain numerical type hyperparameters and categorical type hyperparameters; outlining an axis-parallel box with the numerical type hyperparameters as the edges, and the area covered by the axis-parallel box is the positive sampling area of the numerical type hyperparameters; taking the categorical type hyperparameters as the positive sampling area of the categorical type hyperparameters; Sample in the positive sampling area with a first preset probability and sample in the global sampling area with a second preset probability to select a new hyperparameter group as the hyperparameter group for the next round of training; S105. Optimize the stored hyperparameter groups and select a new hyperparameter group as the hyperparameter group for the next round of training; S106. Repeat steps S102 to S105. When the number of training times or the training duration reaches a preset value, output the optimal training model and its corresponding hyperparameter group.
2. The method for determining hyperparameters of the deep learning model according to claim 1, wherein The first preset probability is greater than the second preset probability.
3. The method for determining hyperparameters of the deep learning model according to claim 1, characterized in that The obtaining a number of hyperparameter groups includes: Select a number of hyperparameter groups for the first evaluation according to the stored experience.
4. The method for determining hyperparameters of the deep learning model according to claim 1, characterized in that The training and evaluating the target model through a number of distributed cluster sub-nodes based on the hyperparameter groups, the training data, and the test data to obtain a number of evaluation values of the training models and hyperparameter groups includes: Allocate the target model, training data, and test data to each of the distributed cluster sub-nodes; Configure one hyperparameter group for each of the distributed cluster sub-nodes; Train and evaluate the target model through a number of the distributed cluster sub-nodes to obtain a number of evaluation values of the training models and hyperparameter groups.
5. An apparatus for determining hyperparameters of a deep learning model, characterized in that, Including: A first obtaining unit for obtaining model data to be trained, where the model data includes a target model, training data, and test data; A second obtaining unit for obtaining a number of hyperparameter groups; A training unit for training and evaluating the target model through a number of distributed cluster sub-nodes based on the hyperparameter groups, the training data, and the test data to obtain a number of evaluation values of the training models and hyperparameter groups; A storage unit for storing the hyperparameter groups and the evaluation values of the hyperparameter groups, and storing the optimal training model from a number of the training models; An optimization unit, configured to optimize the stored hyperparameter groups and select several new hyperparameter groups as the hyperparameter groups for the next round of training; An output unit, configured to output the optimal trained model and its corresponding hyperparameter group when the number of training times or the training duration reaches a preset value.
6. An electronic device, comprising a processor and a memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method for determining hyperparameters of the deep learning model according to any one of claims 1 to 4 are implemented.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the method for determining hyperparameters of the deep learning model according to any one of claims 1 to 4 are implemented.
Citation Information
Patent Citations
Hyper-parameter optimization method and device, computer equipment and storage medium
CN111105040A
Method and system for automatically training machine learning model
CN112085205A