Neural network hyper-parameter optimization method

By combining CMA-ES and TPE, candidate solutions are generated far away from the population region, which solves the problem that the CMA-ES algorithm is prone to fall into local optimality, and achieves the stability and efficiency improvement of neural network hyperparameter optimization.

CN120297370APending Publication Date: 2025-07-11SHANGHAI JIAOTONG UNIV +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510357716.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The candidate solutions generated by the CMA-ES algorithm in the prior art are usually concentrated near the current population mean, and the exploration range is relatively limited, which is easy to fall into local optimization, and it is difficult to effectively solve complex optimization problems.

Method used

Combining the Bayesian optimization algorithm of covariance matrix adaptive evolution strategy (CMA-ES) and tree structure Parson estimator (TPE), we enhance global exploration capabilities by generating candidate solutions far away from the current population area and avoiding early fall into local optimization.

Benefits of technology

It significantly improves the stability and solution efficiency of neural network hyperparameter optimization, and can more efficiently deal with complex optimization problems with multiple local minimums.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120297370A_ABST
    Figure CN120297370A_ABST
Patent Text Reader

Abstract

The invention discloses a neural network hyper-parameter optimization method, which comprises the following steps of: firstly, selecting a search range of a solution space and an upper limit of a population iteration round, and setting an initial value of an iteration population and the number of iteration rounds; then, carrying out iterative solution, generating a plurality of initial candidate solutions through a covariance matrix-based adaptive evolutionary strategy (CMA-ES) in an iterative period, and carrying out fitness evaluation on the initial candidate solutions; sequentially generating a plurality of supplementary candidate solutions through a Bayesian optimization algorithm based on a TPE (Tree Structure Paren Estimator), and carrying out fitness evaluation; synthesizing all candidate solutions of the current round, selecting a plurality of elite individuals with the highest fitness from all the candidate solutions as a next-generation population, updating related parameters of the CMA-ES algorithm, and completing one-time iteration; if the set population iteration round upper limit is not reached, the iteration solving process is repeated, and if the set round upper limit is reached, the optimal solution is output. According to the method, the stability of optimization is remarkably improved, and higher solving efficiency is shown when a complex optimization problem with multiple local minimum values is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a technology in the field of neural networks, and specifically to an optimization method for hyperparameters of neural networks. Background Art

[0002] In recent years, in the hyperparameter optimization problem of complex neural network architectures, evolutionary algorithms have played a key role. Among them, the covariance matrix adaptation evolution strategy (CMA-ES) is regarded as a cutting-edge technology in the field of evolutionary computation and has become one of the standard configurations for continuous optimization in many research laboratories and industrial environments around the world. However, the candidate solutions generated by the CMA-ES algorithm usually concentrate around the current population mean, with a relatively limited exploration range and are prone to falling into local optima, thus requiring further improvement. Summary of the Invention

[0003] Aiming at the deficiencies that the candidate solutions generated by the classic CMA-ES algorithm in the prior art usually concentrate around the current population mean, with a relatively limited exploration range and being prone to falling into local optima, the present invention proposes an optimization method for hyperparameters of neural networks, which significantly improves the stability of optimization and shows higher solution efficiency when dealing with complex optimization problems with multiple local minima.

[0004] The present invention is realized through the following technical solutions:

[0005] The present invention relates to an optimization method for hyperparameters of neural networks. First, the search range of the solution space and the upper limit of the population iteration rounds are selected, and the initial value and the number of iteration rounds of the iterative population are set; then, enter the iterative solution. In one iteration cycle, after generating a number of initial candidate solutions based on the covariance matrix adaptation evolution strategy (CMA-ES) and performing fitness evaluation on them; generate a number of supplementary candidate solutions in sequence through the Bayesian optimization algorithm based on the tree-structured Parzen estimator (TPE) and perform fitness evaluation; then, synthesize all the candidate solutions in the current round, select several elite individuals with the highest fitness as the next generation population, update the relevant parameters of the CMA-ES algorithm, and complete one iteration; if the set upper limit of the population iteration rounds is not reached, repeat the above iterative solution process, and if the set round upper limit is reached, output the optimal solution.

[0006] The generation of a number of initial candidate solutions based on the covariance matrix adaptation evolution strategy (CMA-ES) is specifically as follows: Wherein: is the i-th initial candidate solution in the t-th generation, m t is the mean vector of the t-th generation, σ t is the step size, N i (0, C t) is a random vector generated from a zero-mean, multivariate normal distribution, and C t is the covariance.

[0007] The method of generating a number of supplementary candidate solutions in sequence through a Bayesian optimization algorithm based on a tree-structured Parzen estimator (TPE) is specifically as follows: Where: x is a generated supplementary candidate solution, D is the hyperparameter search space, g(x) = P(x|y≥y * ), is the distribution of solution x when the objective fitness is not lower than a certain threshold, l(x) = P(x|y<y * ), is the distribution of solution x when the objective fitness is lower than a certain threshold, and g(x) and l(x) are constructed using kernel density estimation based on all the solutions and their fitness values generated by the previously mentioned improved covariance matrix adaptation evolution strategy algorithm.

[0008] The fitness evaluation mentioned above refers to obtaining the fitness value by constructing a neural network model based on the hyperparameter combination of the candidate solution and then through the performance of the model on the test set (such as R 2 ).

[0009] Updating the parameters of CMA-ES includes: Where: μ is the number of candidate solutions used to calculate the new mean, and the μ elite individuals with the highest fitness are selected from the n1 initial candidate solutions generated by CMA-ES and the n2 supplementary candidate solutions generated by Bayesian optimization in this optimization round. represents the candidate solution with the i-th best fitness in the t-th generation, ω i is the weight of the i-th candidate solution, c1 and c μ are the learning rates, p c,t+1 is the covariance update path, p σ,t+1 is the step size control path, c σ is the path decay coefficient, d σ is the damping coefficient, and E‖N(0,I)‖ is the expected length of the standard normal distribution. Technical Effects

[0010] Based on the traditional CMA-ES algorithm, the present invention introduces a Bayesian optimization mechanism based on TPE, supplements candidate solutions that may come from areas far from the current population in each generation of the population, provides global exploration information for the optimization process, and thus effectively avoids getting trapped in the local optimum in the early stage. Compared with the prior art, by combining the global exploration ability of the Bayesian optimization algorithm based on TPE with the local development ability of the CMA-ES algorithm, the present invention significantly improves the stability of the optimization and shows higher solution efficiency when dealing with complex optimization problems with multiple local minima. Description of the Drawings

[0011] Figure 1 This is the flow chart of the present invention;

[0012] Figure 2 This is the comparison chart of the results of the embodiments. Detailed implementation manners

[0013] As Figure 1 shown, this embodiment relates to an optimization method for hyperparameters of a neural network, including:

[0014] S1: Select the search range of the solution space and the upper limit of the population iteration rounds, and set the initial value of the iterative population and the number of iteration rounds.

[0015] The search range of the solution space refers to: the learning rate of the neural network model, the L2 regularization coefficient, and the number of neurons in each layer of the hidden layer (in this example, the neural network model uses four hidden layers). Reasonable ranges are taken for these parameters respectively to form the hyperparameter search space.

[0016] The initial value of the iterative population is obtained by random sampling within the search space.

[0017] The upper limit of the iteration rounds is set to N, and the initial population iteration round t = 1 is initialized.

[0018] S2: Generate n1 initial candidate solutions by CMA-ES and evaluate their fitness.

[0019] The initial candidate solutions are hyperparameter configuration vectors, and their elements include the learning rate, regularization coefficient, number of neurons, etc. The initial candidate solutions are specifically: Among them: is the i-th initial candidate solution in the t-th generation, m t is the mean vector of the t-th generation, σ t is the step size, N i (0, C t ) is a random vector generated from a zero-mean, multivariate normal distribution, and C t is the covariance.

[0020] The fitness evaluation refers to: constructing a neural network model using the hyperparameter combination corresponding to the initial candidate solution; training the neural network model using a pre-prepared training set; testing the neural network model using a pre-prepared test set, and the test result is R 2 , which is used as the fitness of this initial candidate solution.

[0021] The R 2 , specifically: Among them: n is the number of test set samples, y i is the test set sample label, is the model prediction value, is the average value of the sample labels in the test set; R 2 less than 1, R 2 The closer it is to 1, the better the prediction performance of the model.

[0022] S3: The Bayesian optimization algorithm based on TPE generates supplementary candidate solutions by maximizing the acquisition function and conducts fitness evaluation, then updates the relevant parameters of the Bayesian optimization algorithm. On this basis, the next supplementary candidate solution is generated. This process is iterated repeatedly until n2 supplementary candidate solutions are generated. The acquisition function is specifically: where: x is a generated supplementary candidate solution, D is the hyperparameter search space, g(x) is the distribution of solutions x when the target fitness y is not lower than the threshold y * and l(x) is the distribution of solutions x when the target fitness y is lower than the threshold y *

[0023] The distributions g(x) and l(x) are estimated by kernel density.

[0024] The fitness evaluation is implemented by the same method as in step S2.

[0025] The number n2 of the supplementary candidate solutions does not have to be the same as the number n1 of the initial candidate solutions.

[0026] S4: Combine all the candidate solutions in the current round, select the μ elite individuals with the highest fitness as the next-generation population, and update the relevant parameters of the CMA-ES algorithm, specifically: where: μ is the number of candidate solutions used to calculate the new mean. The μ elite individuals with the highest fitness are selected from the n1 initial candidate solutions generated by CMA-ES and the n2 supplementary candidate solutions generated by Bayesian optimization in this optimization round. represents the candidate solution with the t-th best fitness in the t-th generation, ω i is the weight of the i-th candidate solution, c1 and c μ are the learning rates, p c,t+1 is the covariance update path, p σ,t+1 is the step-size control path, c σ is the path decay coefficient, d σ is the damping coefficient, and E‖N(0,I)‖ is the expected length of the standard normal distribution.

[0027] S5: Judge whether t < N: If t < N, then go to step S2 to enter the next-generation optimization; otherwise, if the population iteration round reaches the predetermined round upper limit, then go to step S6 to end the optimization, where: N is the population iteration round upper limit set to N.

[0028] ​S6: Output the optimal solution: Output the solution with the highest fitness value during the optimization process, that is, the optimal hyperparameter combination of the neural network model under the training of this data set.

[0029] Through specific actual experiments, a neural network model is trained using the martensite yield strength data set of metal materials. The sample features of this data set are material composition, temperature, ambient pressure, and equivalent plastic strain, and the sample label is the martensite yield strength of the material, with a total of 2311 data. The data set is divided into a test set and a training set. The training set is used for model training, and the test set is used for model testing. The optimization algorithm described in the present invention is used to optimize the hyperparameters of the neural network model, and the relationship between the optimal R 2 with the number of search times is recorded. At the same time, the classical CMA-ES is used to optimize the hyperparameters of this neural network model.

[0030] Table 1 Partial data examples of the martensite yield strength data set of metal materials Material Temperature / °C Ambient pressure / atm Equivalent plastic strain Yield strength / Pa 34CrNi3Mo 300.0 1.0 0.01 1.49602E9 20CrMo 600.0 1.0 1.7 3.506176E8 20CrMnMo 700.0 1.0 0.0075 2.18177E8 20CrMo 1100.0 1.0 2.5 2.01524E7 20Cr2Ni4 200.0 1.0 0.095 1.31146E9 20CrMnTi 400.0 1.0 0.005 9.313361E8 20CrNi2Mo 1100.0 1.0 1.3 2.3006E7

[0031] Table 2 Comparison of optimization effects R2 convergence value Number of convergence steps The present invention 0.99573 105 Traditional CMA-ES algorithm 0.9922 149

[0032] As Figure 2 shown, it is the comparison result of the two algorithms. In the figure, the horizontal axis is the number of search times, and the vertical axis is the optimization objective (R2 score, the closer to 1, the better the model prediction effect). It can be seen from the results that the improved covariance matrix adaptation evolution strategy algorithm proposed in the present invention has higher optimization efficiency than the classical CMA-ES in this problem and achieves better optimization results in a limited number of optimizations. This is because the candidate solutions generated by the classical CMA-ES algorithm usually concentrate near the current population mean, with a relatively limited exploration range and are prone to falling into local optima. While the present invention introduces a Bayesian optimization mechanism based on TPE to supplement candidate solutions in each generation of the population. These solutions may come from areas far from the current population, introducing global exploration information into the optimization process, thereby effectively avoiding the dilemma of falling into local optima in the early stage. By combining the global exploration ability of the Bayesian optimization algorithm based on TPE with the local exploitation ability of the CMA-ES algorithm, the present invention significantly improves the stability of optimization. Compared with the classical CMA-ES algorithm, the present invention shows higher solution efficiency when dealing with complex optimization problems with multiple local minima.

[0033] Compared with the traditional CMA-ES algorithm that is prone to falling into local minima during solution, the improvement of the present invention provides some global information for population iteration, enabling it to jump out of local minima and showing higher solution efficiency when dealing with complex optimization problems with multiple local minima.

[0034] The above specific embodiments can be locally adjusted in different ways by those skilled in the art without departing from the principles and purposes of the present invention. The protection scope of the present invention is subject to the claims and is not limited by the above specific embodiments. All implementation solutions within its scope are subject to the present invention.

Claims

1. An optimization method for neural network hyperparameters, characterized in that, First, select the search range of the solution space and the upper limit of the population iteration rounds, and set the initial value of the iterative population and the number of iteration rounds. Then, enter the iterative solution process. In one iteration cycle, after generating several initial candidate solutions based on the Covariance Matrix Adaptation Evolution Strategy (CMA-ES), perform fitness evaluation on them. Generate several supplementary candidate solutions in sequence through the Bayesian optimization algorithm based on the Tree-structured Parzen Estimator (TPE), and perform fitness evaluation. Then, comprehensively consider all candidate solutions in the current round, select several elite individuals with the highest fitness as the next-generation population, and update the relevant parameters of the CMA-ES algorithm to complete one iteration. If the set upper limit of the population iteration rounds is not reached, repeat the above iterative solution process. If the set round limit is reached, output the optimal solution.

2. The optimization method of neural network hyperparameters according to claim 1, characterized in that Generating a number of initial candidate solutions through the Covariance Matrix Adaptation Evolution Strategy (CMA-ES) specifically as follows: Where: is the i-th initial candidate solution in the t-th generation, m t is the mean vector in the t-th generation, σ t is the step size, N i (0, C t ) is a random vector generated from a zero-mean, multivariate normal distribution, and C t is the covariance.

3. The optimization method for neural network hyperparameters according to claim 1, characterized in that A number of supplementary candidate solutions are successively generated by a Bayesian optimization algorithm based on a tree-structured Parzen estimator (TPE), specifically: maximizing the acquisition function where: x is a generated supplementary candidate solution, D is the hyperparameter search space, g(x) = P(x|y≥y * ), which is the distribution of solution x when the objective fitness is not lower than a certain threshold, l(x) = P(x|y<y * ), which is the distribution of solution x when the objective fitness is lower than a certain threshold, and g(x) and l(x) adopt kernel density estimation.

4. The optimization method of neural network hyperparameters according to claim 1, characterized in that, The fitness evaluation mentioned above refers to constructing a neural network model based on the hyperparameter combination of the candidate solution, and obtaining its fitness value through the performance of the model on the test set (such as R 2 ). That is: constructing a neural network model using the hyperparameter combination corresponding to the initial candidate solution; training the neural network model using the pre-prepared training set; testing the neural network model using the pre-prepared test set, and the test result is R 2 , which is used as the fitness of the initial candidate solution.

5. The optimization method of neural network hyperparameters according to claim 1, characterized in that, The parameters for updating CMA-ES include: Among them: μ is the number of candidate solutions used to calculate the new mean. The μ elite individuals with the highest fitness are selected from the n1 initial candidate solutions generated by CMA-ES and the n2 supplementary candidate solutions generated by Bayesian optimization in this optimization round. represents the candidate solution with the i-th best fitness in the t-th generation, ω i is the weight of the i-th candidate solution, c1 and c μ are learning rates, p c,t+1 is the covariance update path, p σ,t+1 is the step size control path, c σ is the path decay coefficient, d σ is the damping coefficient, and E‖N(0,I)‖ is the expected length of the standard normal distribution.

6. The optimization method of neural network hyperparameters according to any one of claims 1-5, characterized in that specifically Including: S1: Select the search range of the solution space and the upper limit of the population iteration rounds, and set the initial value of the iterative population and the number of iteration rounds; S2: Generate n1 initial candidate solutions by CMA-ES and perform fitness evaluation on them; S3: Generate supplementary candidate solutions through the Bayesian optimization algorithm based on TPE by maximizing the formula, perform fitness evaluation, and update the relevant parameters of the Bayesian optimization algorithm. On this basis, generate the next supplementary candidate solution. This process is iterated repeatedly until n2 supplementary candidate solutions are generated; S4: Comprehensively consider all candidate solutions in the current round, select μ elite individuals with the highest fitness as the next-generation population, and update the relevant parameters of the CMA-ES algorithm; S5: Judge t < N: If t < N, then go to step S2 to enter the next-generation optimization. Otherwise, if the population iteration rounds reach the predetermined round limit, then go to step S6 to end the optimization, where: N is the upper limit of the population iteration rounds set to N; S6: Output the optimal solution: Output the solution with the highest fitness that appears during the optimization process, that is, the optimal hyperparameter combination of the neural network model under the training of this dataset.

Citation Information

Cited By

  • Deep learning assisted acceleration fracturing construction parameter intelligent real-time optimization method

    CN121091692A