Deep Neural Network Hyperparameter Optimization via Surrogate Modeling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for hyperparameter optimization in deep neural networks, such as Bayesian optimization, face challenges with parallelization and dimensionality issues when dealing with a large number of hyperparameters, leading to inefficiencies in finding optimal configurations.

Innovation Solution

A method that samples hyperparameter configurations, trains neural network models, and uses a model-based optimization algorithm with interpolation functions and predictive early stopping to recommend optimal hyperparameter settings, reducing the search space and training time.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If Bayesian optimization algorithm is used for hyperparameter optimization, then automation is improved, but parallelization capability deteriorates and dimensionality issues occur when handling large number of hyperparameters

Engineering Contradiction:
Improvehyperparameter optimization automationVSAvoidparallelization capability
Core Design Contradiction:
Extent of automationVSProductivity

Solution Approach 1:

The patent segments the hyperparameter optimization process into multiple independent parallel tasks. Different worker processes evaluate different hyperparameter configurations simultaneously, dividing the overall optimization workload into parallelizable units that can be executed concurrently across multiple processors or machines.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary manager process that coordinates between the parallel worker processes and the final optimization decision. The manager collects results from parallel workers, updates the surrogate model, and directs the next batch of evaluations, enabling automated optimization while maintaining parallel execution capability.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If comprehensive hyperparameter search is performed, then optimization precision is improved, but training time and computational resources increase significantly

Engineering Contradiction:
Improvehyperparameter optimization precisionVSAvoidtraining time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary actions by training surrogate models (Gaussian processes or neural network surrogates) that approximate the relationship between hyperparameters and model performance. These surrogate models predict outcomes for untested hyperparameter configurations, allowing the system to identify promising candidates without exhaustive trial-and-error training, thus reducing overall training time while maintaining optimization precision.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements feedback loops where actual training results from worker processes are continuously fed back to update the surrogate models. This feedback mechanism refines the predictions over time, enabling the system to converge on optimal hyperparameters more efficiently by learning from previous evaluation outcomes and focusing subsequent searches on high-potential regions of the hyperparameter space.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11537893B2Method and electronic device for selecting deep neural network hyperparameters
Publication Date: 2022.12.27 IND TECH RES INST
  • US11537893B2 patent drawing
  • US11537893B2 patent drawing
  • US11537893B2 patent drawing

AI summary

A method and an electronic device for selecting deep neural network hyperparameters are provided. In an embodiment of the method, a plurality of testing hyperparameter configurations are sampled from a plurality of hyperparameter ranges of a plurality of hyperparameters. A target neural network model is trained by using a training dataset and the plurality of testing hyperparameter configurations, and a plurality of accuracies corresponding to the plurality of testing hyperparameter configurations are obtained after training for preset epochs. A hyperparameter recommendation operation is performed to predict a plurality of final accuracies of the plurality of testing hyperparameter configurations. A recommended hyperparameter configuration corresponding to the final accuracy having a highest predicted value is selected as a hyperparameter setting for continuing training the target neural network model.