AutoML Hardware Model Optimization via Runtime Metrics

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current technologies face challenges in accurately optimizing trained machine learning (ML) hardware models for embedded systems, requiring efficient and cost-effective processors that consume low power, especially for edge applications like smart surveillance and autonomous driving, where selecting the best ML model and processor combination is crucial for performance and efficiency.

Innovation Solution

The system simultaneously and automatically calculates and compares real runtime performance metrics, such as power, performance, and accuracy, for multiple trained AutoML models on various processor hardware chips, eliminating the need to characterize or model the hardware, and optimizing the ML hardware model by running automated models on actual production chips to obtain accurate metrics for selecting an optimized combination.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional ML hardware model optimization methods are used, then hardware characterization and modeling are performed in advance, but this increases device complexity and development time

Engineering Contradiction:
Improveruntime performance metrics accuracyVSAvoidhardware characterization complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent uses actual production hardware chips as direct testbeds instead of creating abstract hardware models or simulations. By running AutoML models directly on the target hardware, the system obtains accurate runtime performance metrics without needing to characterize or model the hardware in advance, thus eliminating the complexity of hardware characterization while maintaining measurement precision.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system performs self-measurement of runtime performance metrics by executing AutoML models on the actual hardware and automatically collecting performance data. This self-service approach eliminates the need for external hardware characterization processes, reducing device complexity while maintaining accurate performance measurement.

Inventive Principle:
Principle #25Self-service

2Measurement precision

If multiple ML models are tested on multiple processors to find the optimal combination, then selection accuracy improves, but testing time and computational resources increase

Engineering Contradiction:
Improvemodel-selection accuracyVSAvoidtesting time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent employs AutoML models that are pre-trained and optimized before being deployed to actual hardware for testing. This preliminary preparation ensures that when models are tested on multiple processors, the testing phase is streamlined and focused, reducing the overall testing time while maintaining accurate model selection. The AutoML framework automatically performs hyperparameter tuning and model optimization in advance.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system automatically varies multiple parameters including different ML model architectures, hyperparameters, and processor configurations simultaneously. By using AutoML to systematically explore the parameter space and automatically identify optimal combinations, the system achieves high model-selection accuracy without manual intervention, thereby reducing the effective testing time required.

Inventive Principle:
Principle #35Parameter changes

3Productivity

If specialized ML hardware processors are developed for edge applications, then performance and efficiency improve, but development cost and complexity increase

Engineering Contradiction:
ImproveML inference efficiencyVSAvoidprocessor specialization complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent demonstrates that existing general-purpose processors can be effectively utilized for ML inference by systematically selecting and optimizing ML models for specific hardware platforms. Instead of developing specialized ML processors, the system achieves high inference efficiency by automatically matching and optimizing ML models for existing processors, thereby maintaining universality while improving productivity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system optimizes ML model parameters and configurations to match the characteristics of existing processors. By automatically adjusting model architecture, hyperparameters, and inference settings based on the target processor's capabilities, the system achieves high inference efficiency on general-purpose hardware without requiring specialized processor development, thus avoiding increased device complexity.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11836589B1Performance metrics of run-time predictions of automated machine learning (AutoML) models run on actual hardware processors
Publication Date: 2023.12.05 MODELCAT INC
  • US11836589B1 patent drawing
  • US11836589B1 patent drawing
  • US11836589B1 patent drawing

AI summary

Systems and methods for optimizing trained ML hardware models by collecting machine learning (ML) training inputs and outputs; selecting a ML model architecture from ML model architectures; training the selected ML model architecture with the ML training inputs and outputs; selecting a hardware processor from hardware processors; and creating a trained ML hardware model by inputting the selected hardware processor with the trained ML model. ML test inputs and outputs, and types of test metrics are selected and used to test the trained ML hardware model to provide runtime test metrics data for ML output predictions made by the trained ML hardware model. The trained ML hardware model is optimized to become an optimized trained ML hardware model using the runtime test metrics by selecting a new selected ML model architecture, selecting a new selected hardware processor, or updating the trained ML model using the runtime metrics test data.