Auto ML Performance Tuning Device for Dynamic Resource Allocation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Automated Machine Learning (Auto ML) for deep learning requires significant computing resources, leading to imbalances in resource allocation for candidate neural network architectures, which can result in inefficient training times and performance.

Innovation Solution

A performance tuning device is integrated with the Auto ML system to optimize resource allocation by using tools like Scikit-Learn, XGBoost, TensorFlow, and Keras, and implementing methods such as neural architecture search via parameter sharing, to dynamically allocate computing resources based on performance index measurements and system resources, determining strategies for single-node or multi-node training.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If significant computing resources are allocated to Auto ML for deep learning, then model performance can be improved, but resource allocation becomes imbalanced and training efficiency decreases

Engineering Contradiction:
Improvemodel performanceVSAvoidtraining efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system dynamically adjusts computing resource allocation based on real-time performance index measurements. The performance tuning device continuously monitors training progress and modifies resource distribution adaptively, transforming the static resource allocation into a dynamic process that responds to actual training needs of different candidate architectures.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system implements a feedback mechanism where performance index measurements from training processes are fed back to the performance tuning device. This feedback loop enables the system to measure actual training performance and adjust resource allocation accordingly, creating a closed-loop control system that optimizes both model performance and training efficiency.

Inventive Principle:
Principle #23Feedback

2Measurement precision

If computing resources are increased for each candidate architecture, then training accuracy improves, but overall system resource utilization becomes unbalanced

Engineering Contradiction:
Improvetraining accuracyVSAvoidresource utilization
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The system applies different computing resource allocation strategies to different candidate architectures based on their specific training characteristics. Instead of uniform resource distribution, the performance tuning device tailors resource allocation to each candidate's needs, optimizing local resource utilization while maintaining overall system balance.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system changes resource allocation parameters dynamically based on performance index measurements. By adjusting computing resource parameters such as GPU memory, CPU cores, and training time limits according to measured training progress, the system optimizes both training accuracy and overall resource utilization efficiency.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11580458B2Method and system for performance tuning and performance tuning device
Publication Date: 2023.02.14 FULIAN PRESION ELECTRONICS (TIANJIN) CO LTD
  • US11580458B2 patent drawing
  • US11580458B2 patent drawing
  • US11580458B2 patent drawing

AI summary

A method for performance tuning in Automated Machine Learning (Auto ML) includes obtaining preset application program interface and system resources of the automatic machine learning system. Performance index measurement values are obtained according to the preset application program interface when the system pre-trains deep learning training model candidates. A distribution strategy and a resource allocation strategy are determined according to the performance index measurement values and the system resources and computing resources of the system are allocated according to the distribution strategy and the resource allocation strategy. The disclosure also provides an electronic device and a non-transitory storage medium.