A hyperparameter optimization method and system based on genetic algorithm and Gaussian process

By adopting a hyperparameter optimization method based on genetic algorithms and Gaussian processes in deep learning, the problems of inefficient hyperparameter search and poor results are solved, and more efficient neural network model training and better performance are achieved.

CN114118372BActive Publication Date: 2025-05-23深圳市广联智通科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111411192.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-25
Publication Date
2025-05-23
Estimated Expiration
2041-11-25

AI Technical Summary

Technical Problem

The search efficiency and poor search results of the existing technology of the super-registration search method are inefficient and have poor search results, resulting in limited training efficiency and effectiveness of deep learning models.

Method used

Using a hyperparameter optimization method based on genetic algorithms and Gaussian processes, the neural network model is trained in parallel and the probability model is fitted using the Gaussian process, and the hyperparameters are selected and adjusted in advance to improve search efficiency and result quality.

Benefits of technology

This method can find a more suitable hyperparameter setting in a short time, improve the training efficiency and effect of neural network models, and solve the efficiency and results of traditional hyperparameter search methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114118372B_ABST
    Figure CN114118372B_ABST
Patent Text Reader

Abstract

The present invention provides a hyperparameter optimization method and system based on genetic algorithm and Gaussian process, which belongs to the field of deep learning and hyperparameter optimization technology, and solves the problems of low search efficiency and poor search results of hyperparameter search methods in the prior art. The present invention includes: 1) preparation of data set; 2) building a neural network model and generating an initial population; 3) parallel training of the neural network model; 4) obtaining the accuracy of the neural network model in the validation set; 5) judging whether the operation end condition is met, and executing step 6 if it is not met; 6) after fitting the probability model based on the Gaussian process, using the acquisition function to select new hyperparameters, and updating the population based on the new hyperparameters, and then going to step 3. The present invention is used for hyperparameter optimization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] A hyperparameter optimization method and system based on genetic algorithm and Gaussian process are used for hyperparameter optimization, belonging to the technical field of deep learning and hyperparameter optimization. Background Art

[0002] Deep learning builds and simulates neural networks that analyze and learn like the human brain, allowing machines to have analytical learning capabilities like humans and recognize data such as text, images, and sounds. The optimization of deep learning network models is one of the difficult challenges faced in practical applications. The training of neural network models is to find the most appropriate network weights, and the hyperparameters selected before training the neural network model will greatly affect the final training results of the model and the time required for training. Hyperparameters are different from general model parameters. Hyperparameters need to be set in advance before training, such as the learning rate. If the learning rate is too small, the convergence speed of the model will be reduced, and if the learning rate is too large, it will make it difficult for the model to converge. Therefore, choosing the right hyperparameters is crucial for the training of network models.

[0003] The purpose of hyperparameter optimization in machine learning is to find the hyperparameters that make the network model perform best on the validation data set. Experienced engineers rely on trial and error to manually optimize hyperparameters, but this method is very dependent on the engineer's experience and knowledge and takes a lot of time. Grid search and random search require a complete training evaluation for each set of selected hyperparameters, which also requires multiple trials and errors. In addition, the previous training results cannot be used to guide the subsequent hyperparameter selection, and the search efficiency is not high. Summary of the invention

[0004] In response to the above research problems, the purpose of the present invention is to provide a hyper-parameter optimization method and system based on genetic algorithm and Gaussian process to solve the problems of low search efficiency and poor search results of hyper-parameter search methods in the prior art.

[0005] In order to achieve the above object, the present invention adopts the following technical solution:

[0006] A hyperparameter optimization method based on genetic algorithm and Gaussian process comprises the following steps:

[0007] Step 1: Prepare the data set, including training set and validation set;

[0008] Step 2: Build a neural network model and generate an initial population based on multiple neural network models, where the hyperparameters of each neural network model are randomly selected in the hyperparameter search space;

[0009] Step 3: Based on the hyperparameters and training set of each neural network model in the population, each neural network model is trained in parallel for one epoch, where one epoch means that the entire training set is fully input into the neural network model for training;

[0010] Step 4: Use the validation set to validate each neural network model after training and obtain the accuracy of each neural network model;

[0011] Step 5: Determine whether the training of each neural network model has reached a given number of iterations. If so, output the weight of the neural network model with the highest accuracy on the validation set, otherwise go to step 6;

[0012] Step 6: After fitting the probability model based on the Gaussian process, use the acquisition function to select new hyperparameters and update the population based on the new hyperparameters. After the update, execute step 3 again.

[0013] Furthermore, the data set is a data set in the field of deep learning, specifically a data set for image recognition or a data set for target detection.

[0014] Furthermore, the neural network model is ResNet-50, and the population size is 5.

[0015] Furthermore, each neural network model in step 3 is distributed on multiple GPUs for parallel training;

[0016] Select SGD with momentum algorithm as the optimizer to update the weights of the neural network model after back propagation during the training process of the neural network model;

[0017] Each neural network model selects cross entropy loss as the loss function, and the formula is as follows:

[0018]

[0019] Among them, m represents the number of types of pictures in the training set, y i Indicates whether the i-th category image is a real label. If it is true, the value is 1, otherwise the value is 0; It represents the predicted value of the neural network model for the i-th category of pictures, and L represents the difference between the true value and the predicted value of the neural network model.

[0020] Further, the specific steps of step 6 are:

[0021] According to the hyperparameters of each neural network model in the population and its accuracy in the validation set in this round, a data set D = {(x 1 ,y 1 ),(x i ,y i )…(x n ,y n )},xi represents the hyperparameters of the i-th neural network model, y i Represents the accuracy of the i-th neural network model, y i =f(x i ), n represents the population size, and f represents the objective function of the optimized x;

[0022] Assume that f obeys Gaussian process and establish a probability model fitted by data set D, that is, f~GP(μ,K), μ is the mean, K is the covariance matrix, and GP refers to Gaussian process; according to the general calculation process of Gaussian process regression, the hyperparameter prediction point x in the hyperparameter search space is * It also obeys the normal distribution, so is the predicted point mean, is the prediction point variance, y * Represents the hyperparameter x at a point in the hyperparameter search space * The accuracy prediction value at, p represents the probability distribution;

[0023] Use the acquisition function to select the optimal hyperparameters:

[0024]

[0025] Among them, y best is the optimal value in the data set D, x represents the hyperparameter at any point in the hyperparameter search space, and y is the predicted accuracy value at x;

[0026] The point in the hyperparameter search space that maximizes the value of the above acquisition function EI(x) is the optimal hyperparameter x for this round. EI ;

[0027] Replace the weight parameters of the neural network model with the lowest accuracy in the population with the weight parameters of the neural network model with the highest accuracy, and set its hyperparameters to x EI , the other neural network models in the population remain unchanged, go to step 3 and train the entire population again.

[0028] A hyperparameter optimization system based on genetic algorithm and Gaussian process includes the following steps:

[0029] Acquisition module: prepare data sets, including training sets and validation sets;

[0030] Building module: Building a neural network model, generating an initial population based on multiple neural network models, where the hyperparameters of each neural network model are randomly selected in the hyperparameter search space;

[0031] Training module: Based on the hyperparameters and training sets of each neural network model in the population, each neural network model is trained in parallel for one epoch, where one epoch means that the entire training set is fully input into the neural network model for training;

[0032] Verification module: Use the verification set to verify each neural network model after training and obtain the accuracy of each neural network model;

[0033] Judgment module: judge whether the training of each neural network model has reached a given number of iterations. If so, output the weight of the neural network model with the highest accuracy on the validation set, otherwise go to step 6;

[0034] Update module: After fitting the probability model based on the Gaussian process, use the acquisition function to select new hyperparameters, and update the population based on the new hyperparameters. After the update, execute step 3.

[0035] Furthermore, the data set in the acquisition module is a data set in the field of deep learning, specifically a data set for image recognition or a data set for target detection.

[0036] Furthermore, the neural network model obtained in the building module is ResNet-50, and the population size is 5.

[0037] Furthermore, each neural network model in the training module is distributed on multiple GPUs for parallel training;

[0038] Select SGD with momentum algorithm as the optimizer to update the weights of the neural network model after back propagation during the training process of the neural network model;

[0039] Each neural network model selects cross entropy loss as the loss function, and the formula is as follows:

[0040]

[0041] Among them, m represents the number of types of pictures in the training set, y i Indicates whether the i-th category image is a real label. If it is true, the value is 1, otherwise the value is 0; It represents the predicted value of the neural network model for the i-th category of pictures, and L represents the difference between the true value and the predicted value of the neural network model.

[0042] Furthermore, the specific implementation logic of the update module is:

[0043] According to the hyperparameters of each neural network model in the population and its accuracy in the validation set in this round, a data set D = {(x 1 ,y 1 ),(x i ,y i )…(x n ,y n )},x i represents the hyperparameters of the i-th neural network model, y iRepresents the accuracy of the i-th neural network model, y i =f(x i ), n represents the population size, and f represents the objective function of the optimized x;

[0044] Assume that f obeys Gaussian process and establish a probability model fitted by data set D, that is, f~GP(μ,K), μ is the mean, K is the covariance matrix, and GP refers to Gaussian process; according to the general calculation process of Gaussian process regression, the hyperparameter prediction point x in the hyperparameter search space is * It also obeys the normal distribution, so is the predicted point mean, is the prediction point variance, y * Represents the hyperparameter x at a point in the hyperparameter search space * The accuracy prediction value at, p represents the probability distribution;

[0045] Use the acquisition function to select the optimal hyperparameters:

[0046]

[0047] Among them, y best is the optimal value in the data set D, x represents the hyperparameter at any point in the hyperparameter search space, and y is the predicted accuracy value at x;

[0048] The point in the hyperparameter search space that maximizes the value of the above acquisition function EI(x) is the optimal hyperparameter x for this round. EI ;

[0049] Replace the weight parameters of the neural network model with the lowest accuracy in the population with the weight parameters of the neural network model with the highest accuracy, and set its hyperparameters to x EI , the other neural network models in the population remain unchanged, go to step 3 and train the entire population again.

[0050] Compared with the prior art, the present invention has the following beneficial effects:

[0051] First, the present invention only trains the neural network model for one epoch, and uses the training results to select and adjust the hyperparameters in advance, which saves a lot of time compared to completing the training of a neural model and then evaluating the hyperparameter results (comprehensive training of a model usually requires multiple epochs, and the present invention selects hyperparameters after each epoch, which saves a lot of time compared to the entire training process). In addition, the neural network models in the population can be trained in parallel without spending more additional training time. In the training process, the previous information is used to adjust the hyperparameters, which is conducive to finding hyperparameters that are more suitable for the current training stage and enhancing the adaptability of the hyperparameter setting.

[0052] 2. The present invention utilizes the accuracy of the trained neural network model on the validation set and establishes a probability model based on the Gaussian process, so that we can use the existing training results to guide the subsequent hyper-parameter selection, which can improve the search efficiency for hyper-parameters and help accelerate the convergence of the network model. The present invention combines the advantages of parallel search and sequence optimization, and solves their shortcomings. While exploring in parallel, other better parameter models can also be used to eliminate bad models. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 It is a flow chart of the present invention.

[0054] Figure 2 Schematic diagram of Gaussian process regression in the invention. DETAILED DESCRIPTION

[0055] The present invention will be further described below in conjunction with the accompanying drawings and specific implementation methods.

[0056] The present invention proposes a hyper-parameter optimization method based on genetic algorithm and Gaussian process, which adjusts the hyper-parameter settings during the training process, builds a probability model based on Gaussian process, and uses the existing training results to guide the hyper-parameter selection, aiming to improve the training effect of neural network model and solve the problems of low search efficiency and poor search results of traditional hyper-parameter search methods.

[0057] The main process of the present invention includes: 1) preparation of data set; 2) building a neural network model and generating an initial population; 3) parallel training of the neural network model; 4) obtaining the accuracy of the neural network model in the validation set; 5) judging whether the running end condition is met, if not, executing step 6; 6) after fitting the probability model based on the Gaussian process, using the acquisition function to select new hyperparameters, and updating the population based on the new hyperparameters, and then going to step 3. The specific implementation steps are as follows:

[0058] 1. Dataset Preparation

[0059] The user needs to collect enough data for training the neural network model, and divide the collected data set into a training set and a validation set in a ratio of 9:1. The training set is used for the parallel training phase of the neural network model and is used for fitting the neural network model. The validation set is used to preliminarily evaluate the generalization ability of the current model, obtain the accuracy of the current model, collect data for the subsequent construction of the data set D, and guide the selection of hyperparameters and the update of the population. The present invention uses the public data set Cifar-10, and divides the original training set of 50,000 pictures into 45,000 training sets and 5,000 validation sets in proportion.

[0060] 2. Build the neural network model structure and generate the initial population

[0061] The present invention uses the PyTorch framework to build a neural network model, uses the mainstream neural network model ResNet-50 as an example, and can also use other neural network models that can achieve corresponding functions. The population size used in the present invention is 5, that is, there are 5 ResNet-50 neural network models in the population, and the hyperparameter settings of each model in the population are different. The initial hyperparameters can be randomly selected for each model in a reasonable hyperparameter search space. The present invention uses the learning rate as a hyperparameter example, that is, the initial learning rate of each neural network model is different to ensure the diversity of individuals in the population.

[0062] 3. Parallel training of neural network models

[0063] The neural network model is trained using the training set divided in step 1. Each neural network model in the population can be distributed on multiple GPUs for parallel training to speed up the training. The present invention selects the SGD of the momentum algorithm as the optimizer. The SGD of the momentum algorithm can reach the extreme point faster with a smaller amplitude. The cross entropy loss is selected as the loss function formula as follows:

[0064]

[0065] Among them, m represents the number of image types, y i Indicates whether it is true that the input is classified into the i-th category of pictures. If so, the value is 1, otherwise it is set to 0; Represents the model's predicted value for the i-th category of images, and L represents the difference between the true value and the predicted value of the neural network model. The weight parameters of the neural network can be continuously updated through the SGD back propagation and gradient descent algorithms of the momentum algorithm to fit the data training set. After one epoch, the training ends and enters the next step. This step is a common neural network model training process.

[0066] 4. Obtain the accuracy of the neural network model on the validation set

[0067] After the neural network model in the population has been trained for one epoch on the training set, the validation set is used as the neural network input to calculate the image accuracy of each network model and evaluate the generalization ability of the current model. It is necessary to record the proportion of the number of accurate predictions of each network model in the population on the validation set to the total number of validation set images, in order to prepare the data set for fitting the probability model in the subsequent steps, and also to provide an evaluation standard for retaining good individuals and eliminating poor individuals in the genetic algorithm.

[0068] 5. Determine whether the operation end conditions are met

[0069] According to the set value of the number of iterations of the neural network model training, the running end condition is reached. If the number of training cycle iterations has reached the specified number, the program ends and outputs the weight parameters of the neural network model that achieves the highest accuracy on the validation set during the entire training process. If the running end condition is not reached, go to step 6 to continue running.

[0070] 6. Fitting the probability model and updating the population

[0071] After fitting the probability model based on the Gaussian process using the hyperparameter settings of the neural network model in the population and its accuracy on the validation set, the acquisition function is used to select the most promising new hyperparameter (optimal new hyperparameter) in the hyperparameter search space, and then go to step 3 for training after updating the population. Specifically:

[0072] According to the hyperparameters of each neural network model in the population and its accuracy in the validation set in this round, a data set D = {(x 1 ,y 1 ),(x i ,y i )…(x n ,y n )},x i represents the hyperparameters of the i-th neural network model, y i Represents the accuracy of the i-th neural network model, y i =f(x i ), n represents the population size, and f represents the objective function of the optimized x;

[0073] Assume that f obeys Gaussian process and establish a probability model fitted by data set D, that is, f~GP(μ,K), μ is the mean, K is the covariance matrix, and GP refers to Gaussian process; according to the general calculation process of Gaussian process regression, the hyperparameter prediction point x in the hyperparameter search space is * It also obeys the normal distribution, so is the predicted point mean, is the prediction point variance, y * Represents the hyperparameter x at a point in the hyperparameter search space * The accuracy prediction value at, p represents the probability distribution;

[0074] Use the acquisition function to select the optimal hyperparameters:

[0075]

[0076] Among them, y best is the optimal value in the data set D, x represents the hyperparameter at any point in the hyperparameter search space, and y is the predicted accuracy value at x;

[0077] The point in the hyperparameter search space that maximizes the value of the above acquisition function EI(x) is the optimal hyperparameter x for this round.EI ;

[0078] Replace the weight parameters of the neural network model with the lowest accuracy in the population with the weight parameters of the neural network model with the highest accuracy, and set its hyperparameters to x EI , the other neural network models in the population remain unchanged, go to step 3 and train the entire population again.

[0079] In summary, the neural network models in the present invention can be trained in parallel, and the performance evaluation and comparison of each neural network model can be used by using an incomplete training process, which can save a lot of time. The probability model based on the Gaussian process helps us select better hyperparameters for the offspring and accelerate the search of the hyperparameter space.

[0080] The above are only representative embodiments of the present invention in many specific application scopes, and do not constitute any limitation on the protection scope of the present invention. Any technical solutions formed by transformation or equivalent replacement fall within the protection scope of the present invention.

Claims

1. A hyperparameter optimization method based on genetic algorithm and Gaussian process, It is characterized in that The steps include: Step 1: Prepare the data set, including training set and validation set; Step 2: Build a neural network model and generate an initial population based on multiple neural network models, where the hyperparameters of each neural network model are randomly selected in the hyperparameter search space; Step 3: Based on the hyperparameters and training set of each neural network model in the population, each neural network model is trained in parallel for one epoch, where one epoch means that the entire training set is fully input into the neural network model for training; Step 4: Use the validation set to validate each neural network model after training and obtain the accuracy of each neural network model; Step 5: Determine whether the training of each neural network model has reached a given number of iterations. If so, output the weight of the neural network model with the highest accuracy on the validation set, otherwise go to step 6; Step 6: After fitting the probability model based on the Gaussian process, use the acquisition function to select new hyperparameters, and update the population based on the new hyperparameters. After the update, execute step 3 again; The specific steps of step 6 are: According to the hyperparameters of each neural network model in the population and its accuracy in the validation set in this round, a data set is constructed. , Indicates The hyperparameters of the neural network model, Indicates The accuracy of the neural network model, , represents the population size, Indicates optimized The objective function of set up Obey the Gaussian process and establish a data set The fitted probability model is , is the mean, is the covariance matrix, Refers to Gaussian process; according to the general calculation process of Gaussian process regression, the hyperparameter prediction point in the hyperparameter search space It also obeys the normal distribution, so , is the predicted point mean, is the prediction point variance, Represents a hyperparameter at a point in the hyperparameter search space The accuracy prediction value at represents a probability distribution; Use the acquisition function to select the optimal hyperparameters: in, It is a dataset The optimal value in represents any point hyperparameter in the hyperparameter search space, for The accuracy prediction value at Use the acquisition function in the hyperparameter search space The point with the largest value is the optimal hyperparameter for this round. ; Replace the weight parameters of the neural network model with the lowest accuracy in the population with the weight parameters of the neural network model with the highest accuracy, and set its hyperparameters to , the other neural network models in the population remain unchanged, go to step 3 and train the entire population again.

2. The hyperparameter optimization method based on genetic algorithm and Gaussian process according to claim 1, It is characterized in that The data set is a data set in the field of deep learning, specifically a data set for image recognition or a data set for target detection.

3. The hyperparameter optimization method based on genetic algorithm and Gaussian process according to claim 2, It is characterized in that The neural network model is ResNet-50, and the population size is 5.

4. The hyperparameter optimization method based on genetic algorithm and Gaussian process according to claim 3, It is characterized in that The neural network models in step 3 are distributed on multiple GPUs for parallel training; Select SGD with momentum algorithm as the optimizer to update the weights of the neural network model after back propagation during the training process of the neural network model; Each neural network model selects cross entropy loss as the loss function, and the formula is as follows: in, Represents the number of types of pictures in the training set, Indicates Whether the class image is a real label, if it is true, the value is 1, otherwise the value is 0; Represents the neural network model for The predicted value of the class image, Indicates the difference between the true value and the predicted value of the neural network model.

5. A hyperparameter optimization system based on genetic algorithm and Gaussian process, It is characterized in that The steps include: Acquisition module: prepare data sets, including training sets and validation sets; Building module: Building a neural network model, generating an initial population based on multiple neural network models, where the hyperparameters of each neural network model are randomly selected in the hyperparameter search space; Training module: Based on the hyperparameters and training sets of each neural network model in the population, each neural network model is trained in parallel for one epoch, where one epoch means that the entire training set is fully input into the neural network model for training; Verification module: Use the verification set to verify each neural network model after training and obtain the accuracy of each neural network model; Judgment module: judge whether the training of each neural network model has reached a given number of iterations. If so, output the weight of the neural network model with the highest accuracy on the validation set, otherwise go to step 6; Update module: After fitting the probability model based on the Gaussian process, use the acquisition function to select new hyperparameters, and update the population based on the new hyperparameters. After the update, execute step 3 again; The specific implementation logic of the update module is: According to the hyperparameters of each neural network model in the population and its accuracy in the validation set in this round, a data set is constructed. , Indicates The hyperparameters of the neural network model, Indicates The accuracy of the neural network model, , represents the population size, Indicates optimized The objective function of set up Obey the Gaussian process and establish a data set The fitted probability model is , is the mean, is the covariance matrix, Refers to Gaussian process; according to the general calculation process of Gaussian process regression, the hyperparameter prediction point in the hyperparameter search space It also obeys the normal distribution, so , is the predicted point mean, is the prediction point variance, Represents a hyperparameter at a point in the hyperparameter search space The accuracy prediction value at represents a probability distribution; Use the acquisition function to select the optimal hyperparameters: in, It is a dataset The optimal value in represents any point hyperparameter in the hyperparameter search space, for The accuracy prediction value at Use the acquisition function in the hyperparameter search space The point with the largest value is the optimal hyperparameter for this round. ; Replace the weight parameters of the neural network model with the lowest accuracy in the population with the weight parameters of the neural network model with the highest accuracy, and set its hyperparameters to , the other neural network models in the population remain unchanged, go to step 3 and train the entire population again.

6. A hyperparameter optimization system based on genetic algorithm and Gaussian process according to claim 5, It is characterized in that The data set in the acquisition module is a data set in the field of deep learning, specifically a data set for image recognition or a data set for target detection.

7. A hyperparameter optimization system based on genetic algorithm and Gaussian process according to claim 6, It is characterized in that The neural network model obtained in the building module is ResNet-50, and the population size is 5.

8. A hyperparameter optimization system based on genetic algorithm and Gaussian process according to claim 7, It is characterized in that Each neural network model in the training module is distributed on multiple GPUs for parallel training; Select SGD with momentum algorithm as the optimizer to update the weights of the neural network model after back propagation during the training process of the neural network model; Each neural network model selects cross entropy loss as the loss function, and the formula is as follows: in, Represents the number of types of pictures in the training set, Indicates Whether the class image is a real label, if it is true, the value is 1, otherwise the value is 0; Represents the neural network model for The predicted value of the class image, Indicates the difference between the true value and the predicted value of the neural network model.

Citation Information

Patent Citations

  • Internet financial credit evaluation method based on PSO-BP neural network

    CN112037012A

  • Default probability prediction method for optimizing gray neural network based on bacterial foraging algorithm

    CN112634019A